Cách in danh sách này vào DataFrame -Python / BeautifulSoup
Oct 22 2020
Đầu ra cho mã này in từng hàng trên trang web được cung cấp bên dưới.
Tuy nhiên nó cũng bao gồm các thẻ. Về cơ bản, tôi muốn in tất cả các hàng thành dataFrame, tôi có thể đặt nó trên Excel.
.text sẽ không hoạt động vì tôi đang sử dụng find_all vì có các thẻ lặp lại trong tên.
Quá trình sẽ như thế nào để loại bỏ các thẻ không mong muốn và sau đó đưa danh sách vào DF, sao chép trang web?
Cảm ơn.
import requests
from bs4 import BeautifulSoup
import pandas as pd
productlinks=[]
r=requests.get(url)
soup= BeautifulSoup(r.content,'html.parser')
content=soup.find_all('tr')
for item in content:
title=item.find_all('td')
print(title)
Trả lời
1 AndrejKesely Oct 22 2020 at 04:03
Cách dễ nhất là sử dụng pandas.read_html:
import pandas as pd
url='https://sitc.sitcancer.org/2020/abstracts/titles/'
df = pd.read_html(url)[0]
print(df)
df.to_csv('data.csv', index=False)
Bản in:
# ... Keywords
0 1 ... Adoptive immunotherapy; Monocyte/Macrophage; T...
1 2 ... CAR T cells; Immune monitoring; Inflammation; ...
2 3 ... Antibody; Biomarkers; Immune monitoring; T cel...
3 4 ... Biomarkers; RNA; Solid tumors; Tumor microenvi...
4 5 ... Antibody; B cell; Biomarkers; Immune monitorin...
.. ... ... ...
730 752 ... Gene expression; Neoantigens; Regulatory T cel...
731 753 ... Gene expression; Neoantigens; Regulatory T cel...
732 754 ... Biomarkers; Chemokine; Chemotherapy; Costimula...
733 755 ... Chemokine; Granulocyte; Myeloid cells; MDSC; T...
734 756 ... Gene expression; Immune contexture; Immune sup...
[735 rows x 6 columns]
Và lưu data.csv(ảnh chụp màn hình từ LibreOffice):