lemmatizza un'intera colonna usando la funzione lambda
Sep 09 2020
Ho testato questo codice per una frase e voglio convertirlo in modo da poter lemmatizzare un'intera colonna in cui ogni riga è composta da parole senza punteggiatura come: deportivas calcetin hombres deportivas shoes
import wordnet, nltk
nltk.download('wordnet')
from nltk.stem import WordNetLemmatizer
from nltk.corpus import wordnet
import pandas as pd
df = pd.read_excel(r'C:\Test2\test.xlsx')
# Init the Wordnet Lemmatizer
lemmatizer = WordNetLemmatizer()
sentence = 'FINAL_KEYWORDS'
def get_wordnet_pos(word):
"""Map POS tag to first character lemmatize() accepts"""
tag = nltk.pos_tag([word])[0][1][0].upper()
tag_dict = {"J": wordnet.ADJ,
"N": wordnet.NOUN,
"V": wordnet.VERB,
"R": wordnet.ADV}
return tag_dict.get(tag, wordnet.NOUN)
#Lemmatize a Sentence with the appropriate POS tag
sentence = "The striped bats are hanging on their feet for best"
print([lemmatizer.lemmatize(w, get_wordnet_pos(w)) for w in nltk.word_tokenize(sentence)])
Supponiamo che il nome della colonna sia df ['keywords'], puoi aiutarmi a usare una funzione lambda per lemmatizzare l'intera colonna come ho lemmatizzato la frase sopra?
Molte grazie in anticipo
Risposte
AvivYaniv Sep 09 2020 at 16:41
Ecco qui:
- Utilizzare
applyper applicare sulle frasi della colonna - Usa l'espressione lambda che ottiene
sentencecome input e applica la funzione che hai scritto, in modo simile a come hai usato nell'istruzione print
Come parole chiave lemmatizzate:
# Lemmatize a Sentence with the appropriate POS tag
df['keywords'] = df['keywords'].apply(lambda sentence: [lemmatizer.lemmatize(w, get_wordnet_pos(w)) for w in nltk.word_tokenize(sentence)])
Come frase lemmatizzata ( joinparole chiave che utilizzano ""):
# Lemmatize a Sentence with the appropriate POS tag
df['keywords'] = df['keywords'].apply(lambda sentence: ' '.join([lemmatizer.lemmatize(w, get_wordnet_pos(w)) for w in nltk.word_tokenize(sentence)]))