So drucken Sie diese Liste in DataFrame -Python / BeautifulSoup

Oct 22 2020

Die Ausgabe für diesen Code druckt jede Zeile auf der unten angegebenen Website.

Es enthält jedoch auch die Tags. Im Wesentlichen möchte ich alle Zeilen in einen dataFrame drucken, den ich in Excel einfügen kann.

.text würde nicht funktionieren, weil ich find_all verwende, da es Tags gibt, die sich im Namen wiederholen.

Wie wäre der Prozess, um die unerwünschten Tags zu entfernen und die Liste dann in einem DF zu haben, der die Website repliziert?

Vielen Dank.

import requests
from bs4 import BeautifulSoup
import pandas as pd
productlinks=[]
r=requests.get(url)
soup= BeautifulSoup(r.content,'html.parser')
content=soup.find_all('tr')
for item in content:
    title=item.find_all('td')
    print(title)

Antworten

1 AndrejKesely Oct 22 2020 at 04:03

Am einfachsten ist es, Folgendes zu verwenden pandas.read_html:

import pandas as pd

url='https://sitc.sitcancer.org/2020/abstracts/titles/'
df = pd.read_html(url)[0]
print(df)
df.to_csv('data.csv', index=False)

Drucke:

       #  ...                                           Keywords
0      1  ...  Adoptive immunotherapy; Monocyte/Macrophage; T...
1      2  ...  CAR T cells; Immune monitoring; Inflammation; ...
2      3  ...  Antibody; Biomarkers; Immune monitoring; T cel...
3      4  ...  Biomarkers; RNA; Solid tumors; Tumor microenvi...
4      5  ...  Antibody; B cell; Biomarkers; Immune monitorin...
..   ...  ...                                                ...
730  752  ...  Gene expression; Neoantigens; Regulatory T cel...
731  753  ...  Gene expression; Neoantigens; Regulatory T cel...
732  754  ...  Biomarkers; Chemokine; Chemotherapy; Costimula...
733  755  ...  Chemokine; Granulocyte; Myeloid cells; MDSC; T...
734  756  ...  Gene expression; Immune contexture; Immune sup...

[735 rows x 6 columns]

Und speichert data.csv(Screenshot von LibreOffice):