Come creare un lineplot min-max per mese?

Sep 26 2020

Dispongo di dati di serie temporali per il conteggio degli annunci di carne bovina al dettaglio e intendo fare in modo che il grafico a linee in pila miri a mostrare su una base media di tre settimane, la quantità di annunci media che i droghieri hanno pubblicato per negozio la scorsa settimana. Per fare ciò, sono riuscito ad aggregare i dati per la stampa e ho provato a creare un grafico a linee che desidero. La motivazione principale si basa sul contesto del problema e sulla trama desiderata . Nel mio tentativo, non sono riuscito a ottenere un grafico a linee molto bello perché non è informativo da capire. Mi chiedo come posso raggiungere questo obiettivo in matplotlib. Qualcuno può suggerirmi cosa dovrei fare dal mio tentativo attuale? qualche idea?

dati riproducibili e tentativo in corso

Ecco i dati riproducibili minimi che ho usato nel mio tentativo attuale:

import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.dates as mdates
import seaborn as sns
from datetime import timedelta, datetime

url = 'https://gist.githubusercontent.com/adamFlyn/96e68902d8f71ad62a4d3cda135507ad/raw/4761264cbd55c81cf003a4219fea6a24740d7ce9/df.csv'

df = pd.read_csv(url, parse_dates=['date'])
df.drop(columns=['Unnamed: 0'], inplace=True)

df_grp = df.groupby(['date', 'retail_item']).agg({'number_of_ads': 'sum'})
df_grp["percentage"] = df_grp.groupby(level=0).apply(lambda x:100 * x / float(x.sum()))
df_grp = df_grp.reset_index(level=[0,1])

for item in df_grp['retail_item'].unique():
    dd = df_grp[df_grp['retail_item'] == item].groupby(['date', 'percentage'])[['number_of_ads']].sum().reset_index(level=[0,1])
    dd['weakly_change'] = dd[['percentage']].rolling(7).mean()
    fig, ax = plt.subplots(figsize=(8, 6), dpi=144)
    sns.lineplot(dd.index, 'weakly_change', data=dd, ax=ax)
    ax.set_xlim(dd.index.min(), dd.index.max())
    ax.xaxis.set_major_formatter(mdates.DateFormatter('%b %Y'))
plt.gcf().autofmt_xdate()
plt.style.use('ggplot')
plt.xticks(rotation=90)
plt.show()

Risultato corrente

ma non sono riuscito a ottenere il grafico a linee corretto che mi aspettavo, voglio riprodurre il grafico da questo sito . È fattibile per raggiungere questo obiettivo? Qualche idea?

trama desiderata

ecco il grafico di esempio desiderato che voglio ricavare da questi dati riproducibili minimi :

Non so come dovrei apportare modifiche al mio attuale tentativo di ottenere la trama desiderata sopra. Qualcuno può sapere un modo possibile per farlo matplotlib? cos'altro dovrei fare? Ogni possibile aiuto sarebbe apprezzato. Grazie

Risposte

5 TrentonMcKinney Sep 26 2020 at 01:07
  • Vedi anche Come creare un grafico min-max per mese con fill_between?
  • Vedere i commenti in linea per i dettagli
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
import calendar

#################################################################
# setup from question
url = 'https://gist.githubusercontent.com/adamFlyn/96e68902d8f71ad62a4d3cda135507ad/raw/4761264cbd55c81cf003a4219fea6a24740d7ce9/df.csv'
df = pd.read_csv(url, parse_dates=['date'])
df.drop(columns=['Unnamed: 0'], inplace=True)
df_grp = df.groupby(['date', 'retail_item']).agg({'number_of_ads': 'sum'})
df_grp["percentage"] = df_grp.groupby(level=0).apply(lambda x:100 * x / float(x.sum()))
df_grp = df_grp.reset_index(level=[0,1])
#################################################################

# create a month map from long to abbreviated calendar names
month_map = dict(zip(calendar.month_name[1:], calendar.month_abbr[1:]))

# update the month column name
df_grp['month'] = df_grp.date.dt.month_name().map(month_map)

# set month as categorical so they are plotted in the correct order
df_grp.month = pd.Categorical(df_grp.month, categories=month_map.values(), ordered=True)

# use groupby to aggregate min mean and max
dfmm = df_grp.groupby(['retail_item', 'month'])['percentage'].agg([max, min, 'mean']).stack().reset_index(level=[2]).rename(columns={'level_2': 'mm', 0: 'vals'}).reset_index()

# create a palette map for line colors
cmap = {'min': 'k', 'max': 'k', 'mean': 'b'}

# iterate through each retail item and plot the corresponding data
for g, d in dfmm.groupby('retail_item'):
    plt.figure(figsize=(7, 4))
    sns.lineplot(x='month', y='vals', hue='mm', data=d, palette=cmap)

    # select only min or max data for fill_between
    y1 = d[d.mm == 'max']
    y2 = d[d.mm == 'min']
    plt.fill_between(x=y1.month, y1=y1.vals, y2=y2.vals, color='gainsboro')
    
    # add lines for specific years
    for year in [2016, 2018, 2020]:
        data = df_grp[(df_grp.date.dt.year == year) & (df_grp.retail_item == g)]
        sns.lineplot(x='month', y='percentage', ci=None, data=data, label=year)
    
    plt.ylim(0, 100)
    plt.margins(0, 0)
    plt.legend(bbox_to_anchor=(1., 1), loc='upper left')
    
    plt.ylabel('Percentage of Ads')
    plt.title(g)
    plt.show()