Cara mencetak daftar ini ke DataFrame -Python / BeautifulSoup
Oct 22 2020
Output untuk kode ini mencetak setiap baris di situs web yang disediakan di bawah ini.
Namun itu juga termasuk tag. Pada dasarnya saya ingin mencetak semua baris menjadi dataFrame, yang dapat saya letakkan di Excel.
. teks tidak akan berfungsi karena saya menggunakan find_all karena ada tag yang berulang dalam nama.
Bagaimana prosesnya untuk menghapus tag yang tidak diinginkan, dan kemudian membuat daftar tersebut menjadi DF, mereplikasi situs web?
Terima kasih.
import requests
from bs4 import BeautifulSoup
import pandas as pd
productlinks=[]
r=requests.get(url)
soup= BeautifulSoup(r.content,'html.parser')
content=soup.find_all('tr')
for item in content:
title=item.find_all('td')
print(title)
Jawaban
1 AndrejKesely Oct 22 2020 at 04:03
Cara termudah adalah dengan menggunakan pandas.read_html:
import pandas as pd
url='https://sitc.sitcancer.org/2020/abstracts/titles/'
df = pd.read_html(url)[0]
print(df)
df.to_csv('data.csv', index=False)
Cetakan:
# ... Keywords
0 1 ... Adoptive immunotherapy; Monocyte/Macrophage; T...
1 2 ... CAR T cells; Immune monitoring; Inflammation; ...
2 3 ... Antibody; Biomarkers; Immune monitoring; T cel...
3 4 ... Biomarkers; RNA; Solid tumors; Tumor microenvi...
4 5 ... Antibody; B cell; Biomarkers; Immune monitorin...
.. ... ... ...
730 752 ... Gene expression; Neoantigens; Regulatory T cel...
731 753 ... Gene expression; Neoantigens; Regulatory T cel...
732 754 ... Biomarkers; Chemokine; Chemotherapy; Costimula...
733 755 ... Chemokine; Granulocyte; Myeloid cells; MDSC; T...
734 756 ... Gene expression; Immune contexture; Immune sup...
[735 rows x 6 columns]
Dan simpan data.csv(tangkapan layar dari LibreOffice):
Taylor Sheridan Baru Menambahkan 1 Bintang 'Yellowstone' Favoritnya ke Pemeran 'Lawmen: Bass Reeves'