Manuel de science des données Python

Mar 31 2023
.
  1. IPython : au-delà du Python normal
  2. Présentation de NumPy
  3. Manipulation de données avec Pandas
  4. Visualisation avec Matplotlib
  1. Premiers pas avec l'apprentissage automatique
  2. Apprendre à partir des données
  3. Régression linéaire
  4. Bayes naïf
  5. k-Voisins les plus proches
  6. Introduction à l'apprentissage automatique avec Scikit-Learn
  7. L'apprentissage automatique en pratique
  8. Le compromis biais-variance
  9. Estimation de la densité du noyau
  10. Analyse des composants principaux
  11. Apprentissage multiple
  12. Regroupement
  13. Arbres de décision et forêts aléatoires
  14. Optimisation basée sur les dégradés
  15. Clustering K-Means
  16. En profondeur : la classification naïve de Bayes
  17. En profondeur : régression linéaire
  18. Approfondissement : Soutenir les machines vectorielles
  19. En profondeur : arbres de décision et forêts aléatoires
  20. Approfondissement : analyse en composantes principales
  21. En profondeur : Apprentissage multiple
  22. Approfondissement : regroupement de k-moyennes
  1. Une visite éclair de Python
  2. L'essentiel du langage Python
  3. IPython : au-delà du Python normal
  4. NumPy
  5. En savoir plus sur le shell système IPython
  6. Matplotlib
  7. SciPy
  8. Scikit-Learn
  9. Apprentissage automatique avec Scikit-Learn
  10. Autres ressources d'apprentissage automatique
  11. Apprentissage automatique pratique : un exemple simple
  12. Bibliographie

import numpy as np

# create a 1D array
a = np.array([0, 1, 2, 3, 4])
print(a)

import pandas as pd

# create a Pandas DataFrame
data = {'name': ['Alice', 'Bob', 'Charlie', 'David'],
        'age': [25, 32, 18, 47],
        'gender': ['F', 'M', 'M', 'M']}
df = pd.DataFrame(data)
print(df)

# filter rows based on a condition
df_filtered = df[df['age'] > 30]
print(df_filtered)

# group data by a column and compute statistics
grouped_data = df.groupby('gender')['age'].mean()
print(grouped_data)

import matplotlib.pyplot as plt
import numpy as np

# create some data to plot
x = np.linspace(0, 10, 100)
y = np.sin(x)

# create a line plot
plt.plot(x, y)
plt.title('Sine Wave')
plt.xlabel('x')
plt.ylabel('y')
plt.show()

# create a scatter plot
x = np.random.randn(100)
y = np.random.randn(100)
plt.scatter(x, y)
plt.title('Random Data')
plt.xlabel('x')
plt.ylabel('y')
plt.show()

from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier

# load the iris dataset
iris = load_iris()

# split the data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(iris['data'], iris['target'], random_state=0)

# create a K-Nearest Neighbors classifier
knn = KNeighborsClassifier(n_neighbors=1)

# fit the classifier to the training data
knn.fit(X_train, y_train)

# predict the classes of the test data
y_pred = knn.predict(X_test)

# compute the accuracy of the classifier
accuracy = knn.score(X_test, y_test)
print('Accuracy:', accuracy)