Cách in danh sách này vào DataFrame -Python / BeautifulSoup

Oct 22 2020

Đầu ra cho mã này in từng hàng trên trang web được cung cấp bên dưới.

Tuy nhiên nó cũng bao gồm các thẻ. Về cơ bản, tôi muốn in tất cả các hàng thành dataFrame, tôi có thể đặt nó trên Excel.

.text sẽ không hoạt động vì tôi đang sử dụng find_all vì có các thẻ lặp lại trong tên.

Quá trình sẽ như thế nào để loại bỏ các thẻ không mong muốn và sau đó đưa danh sách vào DF, sao chép trang web?

Cảm ơn.

import requests
from bs4 import BeautifulSoup
import pandas as pd
productlinks=[]
r=requests.get(url)
soup= BeautifulSoup(r.content,'html.parser')
content=soup.find_all('tr')
for item in content:
    title=item.find_all('td')
    print(title)

Trả lời

1 AndrejKesely Oct 22 2020 at 04:03

Cách dễ nhất là sử dụng pandas.read_html:

import pandas as pd

url='https://sitc.sitcancer.org/2020/abstracts/titles/'
df = pd.read_html(url)[0]
print(df)
df.to_csv('data.csv', index=False)

Bản in:

       #  ...                                           Keywords
0      1  ...  Adoptive immunotherapy; Monocyte/Macrophage; T...
1      2  ...  CAR T cells; Immune monitoring; Inflammation; ...
2      3  ...  Antibody; Biomarkers; Immune monitoring; T cel...
3      4  ...  Biomarkers; RNA; Solid tumors; Tumor microenvi...
4      5  ...  Antibody; B cell; Biomarkers; Immune monitorin...
..   ...  ...                                                ...
730  752  ...  Gene expression; Neoantigens; Regulatory T cel...
731  753  ...  Gene expression; Neoantigens; Regulatory T cel...
732  754  ...  Biomarkers; Chemokine; Chemotherapy; Costimula...
733  755  ...  Chemokine; Granulocyte; Myeloid cells; MDSC; T...
734  756  ...  Gene expression; Immune contexture; Immune sup...

[735 rows x 6 columns]

Và lưu data.csv(ảnh chụp màn hình từ LibreOffice):