このリストをDataFrameに出力する方法-Python / BeautifulSoup
Oct 22 2020
このコードの出力は、以下に提供されているWebサイトの各行を印刷します。
ただし、タグも含まれています。基本的に、すべての行をdataFrameに出力して、Excelに配置したいと思います。
名前が繰り返されるタグがあるため、find_allを使用しているため、.textは機能しません。
不要なタグを削除してから、リストをDFに入れて、Webサイトを複製するプロセスはどのようになりますか?
ありがとう。
import requests
from bs4 import BeautifulSoup
import pandas as pd
productlinks=[]
r=requests.get(url)
soup= BeautifulSoup(r.content,'html.parser')
content=soup.find_all('tr')
for item in content:
title=item.find_all('td')
print(title)
回答
1 AndrejKesely Oct 22 2020 at 04:03
最も簡単な方法は、以下を使用することpandas.read_htmlです。
import pandas as pd
url='https://sitc.sitcancer.org/2020/abstracts/titles/'
df = pd.read_html(url)[0]
print(df)
df.to_csv('data.csv', index=False)
プリント:
# ... Keywords
0 1 ... Adoptive immunotherapy; Monocyte/Macrophage; T...
1 2 ... CAR T cells; Immune monitoring; Inflammation; ...
2 3 ... Antibody; Biomarkers; Immune monitoring; T cel...
3 4 ... Biomarkers; RNA; Solid tumors; Tumor microenvi...
4 5 ... Antibody; B cell; Biomarkers; Immune monitorin...
.. ... ... ...
730 752 ... Gene expression; Neoantigens; Regulatory T cel...
731 753 ... Gene expression; Neoantigens; Regulatory T cel...
732 754 ... Biomarkers; Chemokine; Chemotherapy; Costimula...
733 755 ... Chemokine; Granulocyte; Myeloid cells; MDSC; T...
734 756 ... Gene expression; Immune contexture; Immune sup...
[735 rows x 6 columns]
そして保存data.csv(LibreOfficeからのスクリーンショット):