Kiếm nhà như hacker (Phần I)
Nhiều người trong chúng ta mơ ước mua một ngôi nhà, đây có lẽ sẽ là quyết định tốn kém và quan trọng nhất cho đến nay. Tuy nhiên, không chắc bạn biết thị trường bất động sản để đưa ra quyết định sáng suốt và bạn sẽ dựa vào một số nguồn: thị trường trực tuyến, đại lý bất động sản, bạn bè, quan sát của chính thành phố của bạn, v.v.
Câu hỏi đặt ra: làm thế nào bạn có thể sử dụng các kỹ năng của mình trong một vấn đề thực tế để vượt lên và tăng cơ hội đạt được thỏa thuận tốt ? Chà, thông qua mã hóa và thông tin chi tiết về dữ liệu mạnh mẽ…
Vấn đề là gì?
Thị trường bất động sản cung cấp cho bạn một số phân tích, nhưng nếu đam mê dữ liệu, bạn có thể nhanh chóng nhận ra một số thiếu sót:
- rất nhiều quảng cáo rác: trùng lặp, hết hạn, sai, ngoại lệ, v.v.
- bạn chỉ có một cái nhìn tổng quan về các quảng cáo hiện tại: còn các quảng cáo của ngày hôm qua hoặc tháng trước thì sao, hoặc một quảng cáo bao nhiêu tuổi (nhiều quảng cáo được đăng lại khiến bạn có ấn tượng là chúng mới)? Giá quảng cáo có tăng/giảm không?
- số liệu thống kê vô ích: bạn đang tìm kiếm các quảng cáo cụ thể, trong các khu vực cụ thể, với các đặc điểm cụ thể. Mặc dù bạn có thể lọc những quảng cáo này nhưng bạn không thể điều chỉnh số liệu thống kê mà chúng cung cấp cho bạn.
Bạn cần tận dụng tối đa dữ liệu có sẵn ở đó. Có thêm thông tin sẽ là lợi thế duy nhất của bạn .
Các bước:
- tìm thị trường bất động sản phổ biến nhất (bắt đầu với một)
- xây dựng trình thu thập thông tin, giữ cho dữ liệu của bạn sạch sẽ
- lưu dữ liệu ở định dạng hữu ích
- đừng quá tải chúng với các yêu cầu
- chỉ sử dụng cá nhân, không có gì thương mại
Điều khó khăn nhất là hiểu trang web bạn sắp thu thập thông tin.
Bạn có thể dễ dàng xây dựng trình thu thập thông tin kiểu cũ bằng Python, sử dụng BeautifulSoap . Một ví dụ về các bước mà tôi đã thực hiện:
- bắt đầu với một trang web gốc (ví dụ: quảng cáo được lọc theo thành phố mà bạn quan tâm)
- nhận quảng cáo từng quảng cáo, từng trang
- ánh xạ từng quảng cáo tới đại diện mô hình nội bộ của bạn
- cửa hàng
class Engine:
def __init__(self, website, retrieve, db, pages=1):
seed(int(time.time()))
self.website = website
self.retriever = retrieve
self.pages = pages
self.db = db
def run(self):
page_link = self.website
for page_number in range(0, self.pages):
try:
next_page_link = self.crawl_page(page_link)
print("%s Crawled page %d @ link %s" % (datetime.datetime.now().strftime('%Y-%m-%d %H:%M:%S'), page_number, page_link), flush=True)
page_link = next_page_link
except Exception as e:
print("could not retrieve houses from page %s because of %s." % (page_link, e), flush=True)
def crawl_page(self, page_link):
status_code, text = self.retriever.get(page_link)
if status_code != 200:
raise Exception("Could not retrieve page from page_link %s because we got status code %d. Stopping the program..." % (page_link, status_code))
p = Page(text)
self.insert_houses_from_ad_links(p.get_ad_links())
return p.get_next_page_url()
def insert_houses_from_ad_links(self, ad_links):
shuffle(ad_links)
for ad_link in ad_links:
try:
h = self.get_house_from_ad_link(ad_link)
if h is not None:
self.db.insert_if_not_exists_house(h)
except Exception as e:
print("could not retrieve house from link %s and insert it in the db because of %s" % (ad_link, e), flush=True)
def get_house_from_ad_link(self, ad_link):
status_code, text = self.retriever.get(ad_link, allow_redirects=True)
if status_code != 200:
raise Exception("Could not retrieve ad from ad_link %s because we got status code %d. Skipping ad..." % (ad_link, status_code))
ad = Advertisement(text, ad_link)
chars = ad.get_characteristics()
return House().create(chars)
- quảng cáo trùng lặp, nhưng có cập nhật. Tôi khuyên bạn nên giữ lại tất cả các phiên bản của từng quảng cáo (cung cấp cho chúng một ID) — sau đó bạn có thể sử dụng thông tin để xem quá trình phát triển của quảng cáo
- quảng cáo đã hết hạn (chúng không còn nữa)
class Cleaner:
LIMIT = 100
def __init__(self, retriever, first_page, total_pages, db):
self.retriever = retriever
self.first_page = first_page
self.total_pages = total_pages
self.db = db
def run(self):
page = self.first_page - 1
houses = self.db.get_houses_with_retries(page * self.LIMIT, self.LIMIT, 0, 5)
while len(houses) > 0 and page < self.total_pages:
print("%s Cleaning page %s starting with house %s" % (datetime.datetime.now().strftime('%Y-%m-%d %H:%M:%S'), page+1, houses[0].identifier), flush=True)
for house in houses:
try:
if not house.outdated and self.check_outdated(house):
house.outdated = True
print("Outdated house %s" % house.identifier, flush=True)
self.db.update_house(house)
except Exception as error:
print("could not check outdated for house %s and house url %s because of %s" % (house.identifier, house.url, error), flush=True)
page = page + 1
houses = self.db.get_houses_with_retries(page * self.LIMIT, self.LIMIT, 0, 5)
CREATE DATABASE IF NOT EXISTS imobiliare;
USE imobiliare;
CREATE TABLE IF NOT EXISTS `imobiliare` (
`id` INT(11) AUTO_INCREMENT PRIMARY KEY,
`price` FLOAT(10,2),
`currency` CHAR(3),
`rooms` TINYINT(1),
`build_year` YEAR,
`type_building` VARCHAR(255),
`max_floors` TINYINT(2) unsigned,
`floor` VARCHAR(255),
`comfort` VARCHAR(255),
`layout` VARCHAR(255),
`bathrooms` TINYINT(1) unsigned,
`kitchens` TINYINT(1) unsigned,
`building_structure` VARCHAR(255),
`parking_spots` TINYINT(1) unsigned,
`balconies` TINYINT(1) unsigned,
`height_regime` VARCHAR(255),
`surface_util` FLOAT(7,2),
`surface_util_total` FLOAT(7,2),
`surface_build` FLOAT(7,2),
`city` VARCHAR(255),
`sector` VARCHAR(255),
`district` VARCHAR(255),
`url` TEXT,
`external_id` VARCHAR(255),
`seller` VARCHAR(255),
`geolocation` POINT,
`specification` JSON,
`poi` JSON,
`commission_percentage` float(3,2) DEFAULT NULL,
`status` VARCHAR(255),
`tva_included` BOOLEAN default true,
`outdated` boolean default false,
`publish_date` TIMESTAMP DEFAULT NOW(),
`updated_at` TIMESTAMP DEFAULT NOW() ON UPDATE NOW(),
`created_at` TIMESTAMP DEFAULT NOW()
);
CREATE UNIQUE INDEX idx_publish_date_external_id ON imobiliare (publish_date, external_id);
CREATE INDEX idx_external_id ON imobiliare (external_id);
CREATE INDEX idx_status ON imobiliare (status);
CREATE INDEX idx_outdated ON imobiliare (outdated);
CREATE INDEX idx_updated_at ON imobiliare (updated_at);
CREATE INDEX idx_seller ON imobiliare (seller);
CREATE TABLE IF NOT EXISTS `poligoane` (
`id` INT(11) AUTO_INCREMENT PRIMARY KEY,
`name` VARCHAR(255),
`description` VARCHAR(255),
`colour` VARCHAR(255),
`polygon` Polygon,
`created_at` TIMESTAMP DEFAULT NOW(),
`updated_at` TIMESTAMP DEFAULT NOW() ON UPDATE NOW()
);
CREATE UNIQUE INDEX idx_name ON poligoane(name);
Bạn có thể nói rất cụ thể về những khu vực thành phố mà bạn quan tâm. Bạn có thể sử dụng các đa giác/khu vực thành phố này để lọc quảng cáo của mình, chỉ thu thập dữ liệu những khu vực đó và nhận số liệu thống kê tinh chỉnh. Tôi đã tạo Google Map của riêng mình, nơi bạn có thể sử dụng nhãn.
Bạn có thể nhập dữ liệu từ Google Map
from lxml import etree
import requests
from model.map import Map
def extract_link(config_file_path: str) -> str:
config = etree.parse(config_file_path)
ns = config.getroot().nsmap[None]
link = config.find("x:Document/x:NetworkLink/x:Link/x:href", namespaces={'x': ns})
if link is None:
return ""
return link.text
def get_raw_map(link: str) -> str:
response = requests.get(link, allow_redirects=True)
if response.status_code != 200:
raise ("Could not retrieve page from %s because we got the status code %d." % (link, response.status_code))
return response.text
def get_map(config_file_path: str) -> Map:
link = extract_link(config_file_path)
raw_map_text = get_raw_map(link)
return Map.create_map_from_kml(raw_map_text)
from map.finder import *
from dotenv import load_dotenv
from os import getenv
from infra.db.database import Database
def create_db():
host = getenv("MYSQL_HOST")
db = getenv("MYSQL_DATABASE")
user = getenv("MYSQL_USER")
port = getenv("MYSQL_PORT")
passwd = getenv("MYSQL_PASSWORD")
return Database(host, port, db, user, passwd)
def main():
load_dotenv(dotenv_path=".env", verbose=True)
db = create_db()
m = Mapper(db=db)
m.run()
db.disconnect()
class Mapper:
def __init__(self, db) -> None:
self.db = db
def run(self):
m = get_map(getenv("MAP_CONFIG_RELATIVE_PATH"))
self.db.delete_map_all()
self.db.insert_map_placemarks(m)
if __name__ == '__main__':
main()
from lxml import etree
import keytree
from shapely.geometry import shape, Polygon, Point
class Map:
__create_key = object()
def __init__(self, create_key):
assert (create_key == Map.__create_key), \
"private constructor for Map"
@staticmethod
def create_map_from_kml(raw_map: str):
m = Map(Map.__create_key)
m._parse_map(raw_map)
return m
@staticmethod
def create_map_from_polygons(placemarks):
m = Map(Map.__create_key)
m.placemarks = placemarks
return m
def _parse_map(self, raw_map: str) -> None:
"""
Parses KML file and captures polygons.
KML standard uses long,lat,elevation instead of the natural order lat,long,elevation.
We are reversing the coordinates to the natural order.
:param raw_map:
"""
placemarks = []
m = etree.fromstring(bytes(raw_map, encoding='utf-8'))
ns = m.nsmap[None]
placemarks_raw = m[0].xpath("x:Folder/x:Placemark", namespaces={'x': ns})
for idx, placemark_raw in enumerate(placemarks_raw):
name_raw = placemark_raw.find("x:name", namespaces={'x': ns})
if name_raw is not None:
name = name_raw.text
else:
raise Exception(f'placemark number {idx} is not valid as it has no name')
description = ""
description_raw = placemark_raw.find("x:description", namespaces={'x': ns})
if description_raw is not None:
description = description_raw.text
style_raw = placemark_raw.find("x:styleUrl", namespaces={'x': ns})
if style_raw is not None:
colour = parse_colour(style_raw.text)
else:
raise Exception(f'placemark {name} is not valid as it has no colour')
p = placemark_raw.find("x:Polygon", namespaces={'x': ns})
if p is None:
raise Exception(f'placemark {name} is not valid as it has no polygon')
# reverse coordinates
coord = p.xpath(".//x:coordinates", namespaces={'x': ns})
if len(coord) <= 0:
raise Exception(f'placemark {name} as it has no coordinates in polygon')
for i, c in enumerate(coord):
coord[i].text = reverse_coordinates(c.text)
polygon_shape = shape(keytree.geometry(p))
p = Placemark(name=name, description=description, colour=colour, poly=polygon_shape)
placemarks.append(p)
self.placemarks = placemarks
def contains(self, point: Point) -> list:
valid_polygons = []
[valid_polygons.append(p) for p in self.placemarks if p.polygon.contains(point)]
return valid_polygons
class Placemark:
def __init__(self, name: str, description: str, colour: str, poly: Polygon, identifier=0) -> None:
self.identifier = identifier
self.name = name
self.description = description
self.colour = colour
self.polygon = poly
def contains(self, point: Point) -> bool:
return self.polygon.contains(point)
def print_coordinates(self) -> str:
"""
Prints polygon latitude and longitude in the following format
lat long,
Example:
44.4520689 26.0870725,
44.4468259 26.0966887,
44.4352438 26.1028256
"""
coords = []
for x, y in list(zip(*self.polygon.exterior.coords.xy)):
coords.append('{} {}'.format(x, y))
return ','.join(coords)
Tôi đã chọn có các mô hình riêng biệt và tham gia thông tin theo yêu cầu, thay vì gắn nhãn cho từng quảng cáo khi thu thập thông tin, bởi vì tôi có thể thay đổi Google Map, làm mới dữ liệu của mình và tham gia lại mà không cần thay đổi tất cả hàng triệu quảng cáo mà tôi đã có.
Kết luận đầu tiên
Hiện chúng tôi có hàng triệu điểm dữ liệu về nhà trong cơ sở dữ liệu của mình. Công việc của chúng tôi đã cung cấp cho chúng tôi:
- tất cả quảng cáo (đã hết hạn hay chưa)
- cập nhật quảng cáo
- nhãn phụ
Trong phần tiếp theo, tôi sẽ chỉ cho bạn cách tôi sử dụng dữ liệu này và những số liệu thống kê hữu ích mà tôi có thể nhận được. Chúc mừng

![Dù sao thì một danh sách được liên kết là gì? [Phần 1]](https://post.nghiatu.com/assets/images/m/max/724/1*Xokk6XOjWyIGCBujkJsCzQ.jpeg)



































