헤더를 기반으로 WARC 파일을 청크로 분할 : WARC / 1.0 Python
Oct 06 2020
저는 프로그래밍을 처음 접했고 WARC 파일을 청크로 분할 한 다음 각 청크를 사전에 저장하여 처리하려고합니다.
각 청크는 WARC / 1.0 헤더로 시작해야하며 3 개의 빈 줄로 구분됩니다. 또한 처음 두 단락을 제거하고 싶습니다.
WARC/1.0
WARC-Type: warcinfo
WARC-Date: 2020-08-04T01:43:40Z
WARC-Record-ID: <urn:uuid:959ea654-33fd-466b-b1bf-f08aa8abe774>
Content-Length: 500
Content-Type: application/warc-fields
WARC-Filename: CC-MAIN-20200804014340-20200804044340-00045.warc.gz
isPartOf: CC-MAIN-2020-34
publisher: Common Crawl
description: Wide crawl of the web for August 2020
operator: Common Crawl Admin ([email protected])
hostname: ip-10-67-67-22.ec2.internal
software: Apache Nutch 1.17 (modified, https://github.com/commoncrawl/nutch/)
robots: checked via crawler-commons 1.2-SNAPSHOT (https://github.com/crawler-commons/crawler-commons)
format: WARC File Format 1.1
conformsTo: http://iipc.github.io/warc-specifications/specifications/warc-format/warc-1.1/
# 여기에서 모든 것을 아래로 유지 :
WARC/1.0
WARC-Type: request
WARC-Date: 2020-08-04T03:25:25Z
WARC-Record-ID: <urn:uuid:6c0b749a-4d02-4a77-ab93-9bc4ba094cdc>
Content-Length: 303
Content-Type: application/http; msgtype=request
WARC-Warcinfo-ID: <urn:uuid:959ea654-33fd-466b-b1bf-f08aa8abe774>
WARC-IP-Address: 104.254.66.40
WARC-Target-URI: http://00.auto.sohu.com/d/details?cityCode=450100&planId=1450&trimId=145372
생성기를 사용하여 청크를 그룹화하려고 시도했지만 하나의 그룹 (전체 파일)을 반환합니다. 이를 분리하는 간단한 방법이 있습니까?
라이브러리를 가져올 수 없습니다.
어떤 도움이라도 대단히 감사하겠습니다!
답변
1 GregLindahl Oct 06 2020 at 19:56
이 작업을 수행하는 가장 좋은 방법은 warc 파일을 레코드로 올바르게 구문 분석하는 방법을 알고있는 warcio 라이브러리를 사용하는 것입니다.
그것을 제외하고, 나는 warcio 코드를 당신의 것으로 복사 할 것입니다 (라이센스는 허용됩니다.)
Warc 파일은 복잡하며 완전히 테스트되고 널리 사용되는 라이브러리를 사용하는 것이 올바른 구문 분석 방법입니다.
Common Crawl에서 데이터를 다운로드하는 경우 Python 패키지 cdx_toolkit도 확인하는 것이 좋습니다. 후드 아래에서 warcio를 사용하고 다운로드 단계를 처리합니다.