Cara mencetak daftar ini ke DataFrame -Python / BeautifulSoup

Oct 22 2020

Output untuk kode ini mencetak setiap baris di situs web yang disediakan di bawah ini.

Namun itu juga termasuk tag. Pada dasarnya saya ingin mencetak semua baris menjadi dataFrame, yang dapat saya letakkan di Excel.

. teks tidak akan berfungsi karena saya menggunakan find_all karena ada tag yang berulang dalam nama.

Bagaimana prosesnya untuk menghapus tag yang tidak diinginkan, dan kemudian membuat daftar tersebut menjadi DF, mereplikasi situs web?

Terima kasih.

import requests
from bs4 import BeautifulSoup
import pandas as pd
productlinks=[]
r=requests.get(url)
soup= BeautifulSoup(r.content,'html.parser')
content=soup.find_all('tr')
for item in content:
    title=item.find_all('td')
    print(title)

Jawaban

1 AndrejKesely Oct 22 2020 at 04:03

Cara termudah adalah dengan menggunakan pandas.read_html:

import pandas as pd

url='https://sitc.sitcancer.org/2020/abstracts/titles/'
df = pd.read_html(url)[0]
print(df)
df.to_csv('data.csv', index=False)

Cetakan:

       #  ...                                           Keywords
0      1  ...  Adoptive immunotherapy; Monocyte/Macrophage; T...
1      2  ...  CAR T cells; Immune monitoring; Inflammation; ...
2      3  ...  Antibody; Biomarkers; Immune monitoring; T cel...
3      4  ...  Biomarkers; RNA; Solid tumors; Tumor microenvi...
4      5  ...  Antibody; B cell; Biomarkers; Immune monitorin...
..   ...  ...                                                ...
730  752  ...  Gene expression; Neoantigens; Regulatory T cel...
731  753  ...  Gene expression; Neoantigens; Regulatory T cel...
732  754  ...  Biomarkers; Chemokine; Chemotherapy; Costimula...
733  755  ...  Chemokine; Granulocyte; Myeloid cells; MDSC; T...
734  756  ...  Gene expression; Immune contexture; Immune sup...

[735 rows x 6 columns]

Dan simpan data.csv(tangkapan layar dari LibreOffice):