python yêu cầu lỗi POST, sự cố phiên?

Aug 25 2020

Tôi đang cố gắng bắt chước các hành động sau của trình duyệt thông qua python requests:

  1. Hạ cánh trên https://www.bundesanzeiger.de/pub/en/to_nlp_start
  2. Nhấp vào "Tùy chọn tìm kiếm khác"
  3. Nhấp vào hộp kiểm "Đồng thời tìm dữ liệu đã được lịch sử hóa" (tương ứng với tham số POST isHistorical: true:)
  4. Nhấp vào nút "Tìm kiếm vị thế bán ròng ròng"
  5. Nhấp vào nút "Als CSV herunterladen" để tải xuống tệp csv

Đây là mã tôi phải mô phỏng điều này:

import requests
import re

s = requests.Session()
r = s.get("https://www.bundesanzeiger.de/pub/en/to_nlp_start", verify=False, allow_redirects=True)

matches = re.search(
        r'form class="search-form" id=".*" method="post" action="\.(?P<appendtxt>.*)"',
        r.text
    )
request_url = f"https://www.bundesanzeiger.de/pub/en{matches.group('appendtxt')}"
sr = session.post(request_url, data={'isHistorical': 'true', 'nlp-search-button': 'Search net short positions'}, allow_redirects=True)

Tuy nhiên, mặc dù srcung cấp cho tôi status_code 200, nhưng đó thực sự là một lỗi khi tôi kiểm tra sr.url, nó hiển thịhttps://www.bundesanzeiger.de/pub/en/error-404?9

Tìm hiểu sâu hơn một chút, tôi nhận thấy rằng request_urlở trên giải quyết một cái gì đó như

https://www.bundesanzeiger.de/pub/en/nlp;wwwsid=EFEB15CD4ADC8932A91BA88B561A50E9.web07-pub?0-1.-nlp~filter~form~panel-form

nhưng khi tôi kiểm tra url yêu cầu trong Chrome, nó thực sự

https://www.bundesanzeiger.de/pub/en/nlp?87-1.-nlp~filter~form~panel-form`

87đây dường như thay đổi, cho thấy đó là một số ID phiên, nhưng khi tôi thực hiện việc này bằng cách sử dụng requestsnó dường như không giải quyết đúng cách.

Bất kỳ ý tưởng những gì tôi đang thiếu ở đây?

Trả lời

1 AndrejKesely Aug 25 2020 at 02:28

Bạn có thể thử tập lệnh này để tải xuống tệp CSV:

import requests
from bs4 import BeautifulSoup


url = 'https://www.bundesanzeiger.de/pub/en/to_nlp_start'

data = {
    'fulltext': '',
    'positionsinhaber': '',
    'ermittent': '',
    'isin': '',
    'positionVon': '',
    'positionBis': '',
    'datumVon': '',
    'datumBis': '',
    'isHistorical': 'true',
    'nlp-search-button': 'Search+net+short+positions'
}

headers = {
    'Referer': 'https://www.bundesanzeiger.de/'
}

with requests.session() as s:
    soup = BeautifulSoup(s.get(url).content, 'html.parser')

    action = soup.find('form', action=lambda t: 'nlp~filter~form~panel-for' in t)['action']
    u = 'https://www.bundesanzeiger.de/pub/en' + action.strip('.')    

    soup = BeautifulSoup( s.post(u, data=data, headers=headers).content, 'html.parser' )

    a = soup.select_one('a[title="Download as CSV"]')['href']
    a = 'https://www.bundesanzeiger.de/pub/en' + a.strip('.')    

    print( s.get(a, headers=headers).content.decode('utf-8-sig') ) 

Bản in:

"Positionsinhaber","Emittent","ISIN","Position","Datum"
"Citadel Advisors LLC","LEONI AG","DE0005408884","0,62","2020-08-21"
"AQR Capital Management, LLC","Evotec SE","DE0005664809","1,10","2020-08-21"
"BlackRock Investment Management (UK) Limited","thyssenkrupp AG","DE0007500001","1,50","2020-08-21"
"BlackRock Investment Management (UK) Limited","Deutsche Lufthansa Aktiengesellschaft","DE0008232125","0,75","2020-08-21"
"Citadel Europe LLP","TAG Immobilien AG","DE0008303504","0,70","2020-08-21"
"Davidson Kempner European Partners, LLP","TAG Immobilien AG","DE0008303504","0,36","2020-08-21"
"Maplelane Capital, LLC","VARTA AKTIENGESELLSCHAFT","DE000A0TGJ55","1,15","2020-08-21"


...and so on.
idkhowtocode Aug 24 2020 at 23:56

Nếu bạn kiểm tra https://www.bundesanzeiger.de/robots.txt, trang web này không thích được lập chỉ mục. Trang web có thể đang từ chối quyền truy cập vào tác nhân người dùng mặc định được sử dụng bởi bot. Điều này có thể hữu ích: Yêu cầu Python so với robots.txt