对照项目源码拆解——从配置到入库的完整链路

第 6 讲:追踪一条职位数据的完整生命线

前五讲铺完了理论。这一讲,我们不再"概述模块",而是死死咬住一条数据,看着它从命令行启动到最终变成结构化 JSON 的全过程。每一步都对应到源码的具体行,每一处设计都解释为什么。

我们以企业职位(SpiderCom)的 DOM 抓取路径为主干线。这条路径最长、最复杂,覆盖了系统的大半核心机制。API 直连和 on_response 路径会作为分支穿插讲解。

假设我们现在执行:

python main.py -m cp_full -f 2 -d dev

目标:抓取第 2 号配置分片中所有企业的职位列表,再逐条抓取详情页。


第一阶段:起航 —— 从命令行到一个浏览器窗口

1.1 参数进了谁的口袋

main.py 第 18-28 行,argparse 把命令行参数解析成 args 对象:

args.method   = "cp_full"   # -m
args.file     = "2"         # -f
args.dist     = "dev"       # -d
args.proxy    = ""          # -p(未传,空)
args.pagestart= "2"         # -t(未传,默认 2)
args.company  = ""          # -c(未传,空)

第 315-337 行,__main__ 块只做三件事:

set_logger_debug(args.file)           # ① 日志落盘到 log/log_2.txt
s  = SpiderSch(args.file)             # ② 学校采集器(备用)
cs = SpiderCom(args.file)             # ③ 企业采集器(主角)
d  = SpiderData(s)                    # ④ 数据处理器(共享 s 的配置)
run_periodically(s, cs, d)            # ⑤ 进入调度

为什么 SpiderData 要接收 SpiderSch 而不是 SpiderCom 因为 SpiderData 需要访问 INI 配置、浏览器路径、工具路径——这些都在 SpiderSch 的方法里。两个采集器共享同一套基础设施,没必要给 SpiderData 传两个参数。

1.2 run_periodically 的路由

第 151-177 行,run_periodically 先把参数打包进 _stat 字典:

_stat['method']     = "cp_full"
_stat['page_start'] = 2
_stat['dist']       = "dev"
_stat['retry']      = "1"
_stat['p_count']    = 0
_stat['all_proc_list'] = []

这个 _stat 字典会贯穿整个调用链——它不是全局变量,而是作为参数逐层传递。用字典而不是对象的原因:字典可以随时加字段,不需要改任何函数签名。_stat['method']_stat['total']_stat['pfile'] 这些字段是不同函数在不同阶段加进去的。

然后命中第 175-178 行的分支:

if args.method in ["cp", "cp_full"]:
    _stat['method'] = args.method
    clawler_main(cs, _stat)

1.3 clawler_main:浏览器的一生

第 32-71 行。这个函数的名字暴露了它的本质——它就是爬虫的"生命周期管理器"。

def clawler_main(s, _stat=None):
    executable_path = s.get_browser_path()    # ① 拿浏览器路径
    with sync_playwright() as p:              # ② 启动 Playwright
        browser = get_browser(p, executable_path, args.proxy)  # ③ 创建浏览器实例
        s.browser = browser                   # ④ 注入到爬虫对象
        page = browser.new_page()             # ⑤ 开一个标签页

        nodes = s.get_nodes()                 # ⑥ 读配置,拿到所有要爬的公司
        process = s.get_progress()            # ⑦ 读进度文件,知道从哪开始

        for _key, _node in nodes.items():     # ⑧ 逐个公司处理
            for _sch_info in _node:
                if _key in process:
                    s.run(page, _key, _sch_info, _stat)  # ⑨ 核心

        time.sleep(10)
        browser.close()                       # ⑩ 关闭浏览器

几个关键设计:

with sync_playwright() as p:上下文管理器保证 Playwright 进程一定被释放。如果不用 with,程序中途崩了 Playwright 进程会留在后台。

s.browser = browser:把浏览器实例挂在爬虫对象上,这样 SpiderCom 的任何方法都能通过 self.browser 拿到它——比如 get_page_detail_content 里要新开 tab,就需要 self.browser.new_page()

process = s.get_progress():进度文件的用法——不是读"已完成列表",而是读"最后完成到哪了"。后面在 run() 的最后会 write_process_file(_key) 更新这个进度。


第二阶段:落位 —— SpiderCom 如何"认识"一家公司

2.1 __init__:三层配置叠加

spider_com.py 第 42-64 行:

def __init__(self, _file="99"):
    self.config = configparser.ConfigParser()
    self.config.read("data/setting_default.ini", encoding="utf-8")   # 第一层
    self.config.read("data/setting_template.ini", encoding="utf-8")  # 第二层
    self.config.read(f"data/setting_com_{_file}.ini", encoding="utf-8")  # 第三层

Python 的 configparser.read()增量式的——同一个 [section] 下的同一个 key,后读的覆盖先读的。所以:

  • setting_default.ini 定义全局默认值(浏览器路径、保存目录、通用超时时间)
  • setting_template.ini 定义站点模板("百度系职位列表"、"兴业系详情页" 等)
  • setting_com_2.ini 定义第 2 号分片具体包含哪些公司

这就是第 5 讲讲的"default → override"模式的落地实现——不需要任何继承或组合的代码,configparser 自带覆盖语义。

2.2 进度文件:一行文字决定从哪开始

第 49-57 行:

self.progress_list = []
if os.path.exists(self.get_progress_file()):
    with open(self.get_progress_file(), "r", encoding="utf-8") as f:
        for line in f:
            line = line.strip()
            if line:
                self.progress_list = [line]  # 只取第一行
                break

data/progress_com_2.txt 的内容可能就一行:com_00005。意思是"上次跑到了 com_00005,这次从它之后继续"。

2.3 get_nodes():把 INI 变成可迭代的公司列表

第 149-161 行:

def get_nodes(self):
    nodes = {}
    keys = [k for k in self.config.options("Company") if k.startswith("com_")]
    for _key in sorted(keys):
        _svalue = self.config.get("Company", _key)
        _value = json.loads(_svalue)
        self.supplement_node_info(_value)   # 注入模板字段
        nodes[_key] = _value
    return nodes

setting_com_2.ini 里大概是这样的:

[Company]
com_00001 = [{"com_name":"百度","com_webname":"Baidu","template":"tpl_baidu","data_proc_type":"api",...}]
com_00002 = [{"com_name":"京东","com_webname":"JD","template":"tpl_jd","data_proc_type":"api",...}]
com_00003 = [{"com_name":"某某公司","com_webname":"XX","template":"tpl_dom",...}]

注意 template 字段。第 175-184 行的 supplement_node_info 会把模板里的字段注入但不会覆盖

def supplement_node_info(self, _node):
    for _com_info in _node:
        _template = _com_info.get("template")
        if _template:
            _tv = self.config.get("Template", _template)
            _tvjson = json.loads(_tv)
            for _key, _va in _tvjson.items():
                if not _key in _com_info:       # ← 关键:只在公司没定义时才注入
                    _com_info[_key] = _va

if not _key in _com_info——公司自身配置的优先级永远高于模板。例如模板定义了 table_selector: ".job-list",但某家公司用的是 .career-list,公司在自己的配置里写一个 table_selector 就能覆盖模板。


第三阶段:采集 —— 从列表页 HTML 到一条条详情文件

现在浏览器打开了,nodes 拿到了,process 告诉我们要从哪个公司开始。clawler_main 的循环进入了 s.run(page, "com_00003", com_info, _stat)

3.1 run():三通道分叉口

spider_com.py 第 406-473 行。这个函数的第一件事不是打开页面,而是判断用哪种采集方式

def run(self, page, _key, com_info, _stat):
    data_proc_type = com_info.get("data_proc_type", "")
    if data_proc_type == "api":
        return self.api_proc(page, _key, com_info, _stat)
    elif data_proc_type == "on_response":
        return self.on_resp_proc(page, _key, com_info, _stat)

    # 以下:默认的 DOM 抓取路径
    urls = com_info.get("urls")
    for i, k in enumerate(urls):
        url = urls.get(k)
        # ...

三条路的分叉取决于 INI 配置文件里 data_proc_type 这一个字段:

data_proc_type 走的方法 适用场景
"api" api_procauto_api/baidu_data_proc_api.py 有公开职位接口的大厂(百度、京东、金蝶)
"on_response" on_resp_procauto_on_response/main_proc.py 接口有签名、但浏览器能正常调的站点(兴业银行)
空或其他 DOM 抓取 没有接口可用的普通企业站

我们先走 DOM 主干道。

3.2 打开列表页

# 第 425-438 行
pre_open_url = com_info.get("pre_open_url")
if pre_open_url:
    _ok = self.open_with_url(page, pre_open_url)   # 先"预热"首页
    time.sleep(FIRST_PAUSE_TIME)

_ok = self.open_with_url(page, url)                 # 再打开列表页

pre_open_url 的设计是踩坑踩出来的。很多企业招聘系统要求先访问首页建立 Session/Cookie,再访问列表页才给数据。没有这一步,列表页直接返回 403 或空白。

open_with_url(第 240-269 行)做了四层保障:

def open_with_url(self, page, url, refer=""):
    response = page.goto(url, timeout=PAGE_TIMEOUT)
    if response:
        status = response.status
        if status in [200, 412]:                    # 兰州大学返回 412 的特殊兼容
            page.wait_for_load_state('load')         # 等 DOM 加载完
            try:
                page.wait_for_load_state('networkidle', timeout=30000)  # 等网络安静
            except:
                pass                                 # networkidle 超时不算失败
            time.sleep(3)
            return True
    elif page.url == url:                            # 无 response 对象但 URL 确实变了
        return True                                  # (某些 SPA 站点的情况)

networkidle 超时被 except 吞掉了——因为很多页面有持续的心跳请求或 WebSocket,永远不会真正 idle。等 30 秒足够 JS 渲染完,超时就超时,数据已经有了。

3.3 定位列表容器——选择器回退机制

第 482-486 行:

_ok, table_selector = self.get_selector_text(
    page, sch_info, "table_selector", "table_selectors"
)
if not _ok:
    ner_logger.error(f"列表页面没有找到元素:{table_selector},人工处理!")
    return _ret_list

get_selector_text(第 193-238 行)的完整逻辑:

def get_selector_text(self, page, sch_info, selector1, selector2, style3=""):
    # 第一步:用主选择器
    table_selector = sch_info.get("table_selector")  # 如 ".job-list-container"
    style_element = page.query_selector(table_selector)

    if not style_element:
        # 第二步:备选选择器列表(用 | 分隔)
        table_selectors = sch_info.get("table_selectors")  # 如 ".list|.career-list|.position-wrap"
        if table_selectors:
            for _selector in table_selectors.split("|"):
                style_element = page.query_selector(_selector)
                if style_element:
                    table_selector = _selector
                    break

    if not style_element:
        # 第三步:正则匹配 class 名
        selector1_re = sch_info.get("table_selector_re")  # 如 "job.*list"
        if selector1_re:
            div_elements = page.query_selector_all('div')
            for element in div_elements:
                class_name = element.get_attribute('class') or ''
                if re.search(regex_pattern, class_name):
                    return True, f"div.{class_name}"

三层回退:精确选择器 → 备选列表逐个试 → 正则模糊匹配。每一层覆盖一种改版场景:

  • 精确选择器失效:站点小改,换了 class 名 → 备选列表兜底
  • 备选全部失效:站点大改,class 命名规则都变了 → 正则模糊匹配(如 job.*list 能匹配 jobNewListjob_2024_list

为什么正则匹配只查 div 标签? 因为列表容器 99% 是 div。查所有标签太慢,而且误匹配率高。

3.4 动态加载:滚动 + 点击"更多"

第 489-490 行,在选择器定位成功之后:

self._auto_scroll_to_bottom(page)    # 触发懒加载
self._click_load_more(page, sch_info) # 点击"加载更多"按钮

自动滚动

第 300-314 行:

def _auto_scroll_to_bottom(self, page, *, max_scrolls=9999, sleep_s=2.0):
    last_height = page.evaluate("document.body.scrollHeight")
    scroll_count = 0
    while scroll_count < max_scrolls:
        page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
        time.sleep(sleep_s)
        new_height = page.evaluate("document.body.scrollHeight")
        if new_height == last_height:
            return                           # 高度稳定 = 没新内容了
        last_height = new_height
        scroll_count += 1

判停条件是"页面高度不变"而不是"滚了多少次"。一次滚到底 vs 分十次渐进加载,都能正确处理。

点击加载更多

第 316-358 行。支持两种模式:

if load_more_method == "element":
    # 模式 A:配置 CSS 选择器,直接点击
    load_more_button = page.query_selector(load_more_selector)
    if load_more_button and load_more_button.is_visible():
        load_more_button.click()
else:
    # 模式 B:调用站点专属函数(处理复杂的点击逻辑)
    self.pre_page_run(page, sch_info, "click_load_more_func_name")

模式 B 通过 pre_page_run(第 361-372 行)动态导入并执行:

def pre_page_run(self, page, sch_info, func_name="table_func_name"):
    table_func_name = sch_info.get(func_name)         # 如 "click_00088"
    if table_func_name:
        package_func_name = f"auto_gen_com.gen.{table_func_name}"
        return execute_page_action(package_func_name, page)

execute_page_actionauto_gen/func_call.py 第 52-61 行:

def execute_page_action(module_name, page):
    module = importlib.import_module(module_name)
    func = getattr(module, "crawl_page")
    func(page)

所以一个"点击加载更多"可以是任意复杂的 Playwright 操作——先滚动到按钮位置、等它出现、关掉弹窗、再点击——全部封装在一个 auto_gen_com/gen/click_xxxxx.py 文件里。

3.5 提取列表 HTML → 落盘 → 调用解析函数

第 492-513 行。列表容器的 HTML 拿到了,接下来:

tableObj = page.locator(table_selector)
outtext = ["<div>", tableObj.inner_html(), "</div>"]

# ① HTML 先落盘(数据安全)
key_tmp_dir = self.get_key_dir(_key)
_hash = hashlib.md5(url.encode("utf-8")).hexdigest()
tmp_file = f"{key_tmp_dir}/index_{_hash}.html"
with open(tmp_file, "w", encoding="utf-8") as f:
    f.write("\n".join(outtext))

# ② 调用站点专属解析函数:HTML → JSON 列表
func_name = sch_info.get("func_name")                    # 如 "gen_00010"
package_func_name = f"auto_gen_com.gen.{func_name}"
tmp_fname = f"{key_tmp_dir}/index_{_hash}.json"
_ok = call_func(package_func_name, _context_outtext, tmp_fname)

call_funcauto_gen/func_call.py 第 44-49 行:

def call_func(func_name, html_content, tmp_fname):
    _ok = load_and_execute(func_name, 'extract_table_from_html', html_content, tmp_fname)
    return _ok

动态导入 auto_gen_com/gen/gen_00010.py,调用其中的 extract_table_from_html 函数。一个典型的解析函数长这样(gen_00001.py):

def extract_table_from_html(htmlcontext, tempfile):
    soup = BeautifulSoup(htmlcontext, 'html.parser')
    result_list = []
    info_lists = soup.find_all('ul', class_='infoList')
    for ul in info_lists:
        name_tag = ul.find('li', class_='span7')
        time_tag = ul.find('li', class_='span4')
        if name_tag and name_tag.a:
            result_list.append({
                "announcement_name": name_tag.a.text.strip(),
                "publish_time": time_tag.text.strip() if time_tag else "",
                "link": name_tag.a['href'].strip()
            })
    with open(tempfile, 'w', encoding='utf-8') as f:
        json.dump(result_list, f, ensure_ascii=False, indent=4)

如果这个函数返回空列表怎么办? func_call.py 第 26-42 行的 check_result 会触发自动修复:

def check_result(module_name, html_content, tmp_file):
    with open(tmp_file, 'r', encoding='utf-8') as f:
        result_json = json.loads(f.read())
        if len(result_json) == 0:
            gen_func_bygpt(module_name, html_content)   # ← 调 LLM 自动生成新解析函数

gen_func_bygpt 把 HTML 喂给大模型,让它生成新的 extract_table_from_html 函数代码,写入 data/gen_func_code_*_tmp.py。人工审核后改名放到 auto_gen_com/gen/ 目录,下次就能用了。

这一步完成时,磁盘上多了一个 index_*.json 文件,里面是列表页所有职位的数组:

[
  {"announcement_name": "Java开发工程师", "publish_time": "2026-05-01", "link": "/job/12345"},
  {"announcement_name": "产品经理", "publish_time": "2026-05-02", "link": "/job/12346"},
  ...
]

3.6 逐条进入详情页

第 520-527 行,加载 index_*.json 后逐条处理:

with open(tmp_fname, "r", encoding="utf-8") as f:
    _data = json.load(f)
    for i, _item in enumerate(_data):
        _ok = self.get_page_detail_data(page, _key, url, k, key_tmp_dir, sch_info, _item)
        time.sleep(get_random_number())

get_page_detail_data 是整条数据生命线上最复杂的函数(第 531-730 行,近 200 行)。我们一步步拆。

3.6.1 链接的四种形态及处理

第 550-595 行:

_link = _item.get("link")

# 形态 1:void(0) → 链接没法直接用
if _link and _link.startswith(""):
    _link = ""

# 形态 2:没有链接 → 通过点击标题文字获取
if not _link or _click_text == 'Y':
    _text = _item.get("announcement_name")
    # 先查缓存
    _hash = hashlib.md5(combined_text.encode("utf-8")).hexdigest()
    tmp_file = f"{key_tmp_dir}/detail_{_hash}.url"
    if os.path.exists(tmp_file):
        with open(tmp_file, "r") as f:
            _link = f.read()         # ← 从缓存读

    if not _link or _click_text == 'Y':
        # 用 Playwright 在页面上点击标题文字
        new_url, content = click_by_text_and_get_url(
            page, url, _text, _click_type, area, _max_parent_level, _current_url
        )
        if new_url:
            _link = new_url
            with open(tmp_file, "w") as f:
                f.write(_link)       # ← 缓存起来

.url 缓存文件的价值:通过点击获取链接需要操作 DOM、等弹窗/新 tab 打开——慢且不稳定。但同一个标题对应的链接不会变。第一次点击后缓存,下次重跑直接从文件读。

3.6.2 URL 拼接与前端路由保护

第 596-609 行:

domain = com_info.get("json_domain")
_fullurl = self.get_full_url(domain, _link)

# 前端路由保护:带 # 的 URL 不能跟 HTTP 重定向
if "#" not in _fullurl:
    _final_link = get_final_url(_fullurl)
    if _fullurl != _final_link:
        _fullurl = _final_link
else:
    ner_logger.debug(f"检测到前端路由URL,直接使用原始URL: {_fullurl}")

get_final_url 会跟踪 HTTP 重定向。但 https://example.com/#/job/12345 中的 #/job/12345 是前端路由——它不发往服务器。如果跟踪了重定向,可能变成 https://example.com/login(未登录跳转),链接就丢了。

3.6.3 缓存命中:不重复抓取

第 616-653 行:

_hash = hashlib.md5(_fullurl.encode("utf-8")).hexdigest()
tmp_file = f"{key_tmp_dir}/detail_{_hash}.html"
tmp_json_file = f"{key_tmp_dir}/detail_{_hash}.json"

if os.path.exists(tmp_file) and os.path.exists(tmp_json_file):
    # 文件已存在 → 只更新修改时间,不重新抓取
    current_time = time.time()
    os.utime(tmp_file, (current_time, current_time))
    os.utime(tmp_json_file, (current_time, current_time))
    return True

只更新 mtime 不重新抓——这条数据可能在之前的某次运行中已经抓过了。更新 mtime 是为了防止它被"10 天过期"逻辑误清理。

3.6.4 打开详情页:新开 tab 的隔离策略

第 658-676 行——没有命中缓存,进入真正的抓取:

if _context_outtext == "":
    _ok, _context_outtext = self.get_page_detail_content(page, sch_info, domain, _fullurl)

get_page_detail_content 第 733-825 行。和 SpiderSch 在当前页面直接跳转不同,SpiderCom 新开 tab

def get_page_detail_content(self, page, sch_info, domain, _fullurl, _redirect=True):
    page = self.browser.new_page()     # ← 关键:新开 tab
    try:
        response = page.goto(_fullurl, wait_until="networkidle", timeout=3200000)
        page.wait_for_timeout(1500)
    except Exception as e:
        page.close()
        return False, ""

    # 尝试关掉 cookie/同意弹窗
    for btn in ["Accept", "同意", "Continue", "OK"]:
        try:
            page.locator(f"button:has-text('{btn}')").click(timeout=1500)
        except:
            pass

为什么新开 tab? 企业招聘网站的详情页经常注入大量 JS 状态(全局 store、路由状态等),在当前页面跳转后返回列表页,可能触发重新加载或丢失滚动位置。新开 tab 完全隔离了两个上下文。

弹窗关闭的 try/except 全吞异常——因为弹窗不一定存在,而且每种弹窗的语言、文案都可能不同。尝试关但不强求。

3.6.5 正文提取:选择器 → 正则 → 全文

第 762-778 行:

_ok, detail_selector = self.get_selector_text(
    page, sch_info, "detail_selector", "detail_selectors"
)
if _ok:
    detailObj = page.locator(detail_selector)
    if detailObj.count() == 1:
        _context_outtext = f"<div>{detailObj.inner_html()}</div>"
        page.close()
        return True, _context_outtext

# 没找到指定选择器 → 全文返回
page.close()
return True, page.content()

正文提取也是三层回退:指定选择器 → 备选列表 → 正则匹配 → 全文兜底。注意这里的 detailObj.count() == 1——如果匹配到多个元素就不确定哪个是正文,宁可返回全文让后续的大模型来处理,也不乱选一个。

3.6.6 落盘:detail_*.html + detail_*.json

第 680-695 行,回到 get_page_detail_data

# 保存 HTML
with open(tmp_file, "w", encoding="utf-8") as f:
    f.write(_context_outtext)

# 保存 JSON(列表字段 + 详情页元信息)
with open(tmp_json_file, "w", encoding="utf-8") as f:
    _item['full_url'] = _fullurl
    _item['last_url'] = _last_url
    _item['file_path'] = tmp_file
    _item['parent_url'] = url
    _item['channel'] = _key
    _item['job_type'] = k.split("_")[0]      # "shezhao" / "xiaozhao" / "shixi"
    json.dump(_item, f, ensure_ascii=False)

此时,磁盘上多了一对文件

data/tmp/com_00003/
├── index_a1b2c3.html          ← 列表页 HTML 片段
├── index_a1b2c3.json          ← 列表解析结果
├── detail_d4e5f6.html         ← 职位详情 HTML
├── detail_d4e5f6.json         ← 职位元信息
├── detail_d4e5f6.url          ← (可选)点击获取的 URL 缓存
├── detail_g7h8i9.html         ← 第二个职位
├── detail_g7h8i9.json
...

detail_*.json 的内容大概是:

{
  "announcement_name": "Java开发工程师",
  "publish_time": "2026-05-01",
  "link": "/job/12345",
  "full_url": "https://example.com/job/12345",
  "last_url": "https://example.com/job/12345",
  "file_path": "data/tmp/com_00003/detail_d4e5f6.html",
  "parent_url": "https://example.com/careers",
  "channel": "com_00003",
  "job_type": "shezhao"
}

分支 A:API 直连路径(以百度为例)

如果 data_proc_type == "api",则走这条路径。对百度而言,O(浏览器操作) 全部被 O(HTTP POST) 替代。

A.1 入口

spider_com.py 第 374-378 行:

def api_proc(self, page, _key, com_info, _stat):
    return self._special_proc(page, _key, com_info, _stat, label="api", proc=auto_api_proc)

_special_proc 第 380-404 行:循环 urls,每次调 auto_api_proc

A.2 公司路由

auto_api/baidu_data_proc_api.py 第 155-171 行:

def api_proc(spider_com, _key, com_info, k, url, _stat):
    if _key == "com_90001":
        api_proc_baidu(...)
    if _key == "com_90002":
        api_proc_isoftstone(...)
    if _key == "com_90003":
        api_proc_jd(...)
    # ...

为什么用 if-elif 而不是策略模式? 因为每家公司的 API——请求参数名、返回字段名、分页方式、认证方式——完全不同。百度是 recruitType + pageSize + curPage,京东可能完全不一样。强行抽象出一个"统一 API 适配器接口"只会得到一个永远只有一个实现的接口。

A.3 百度:分页 POST → 直接拿 JSON

第 173-245 行:

def api_proc_baidu(spider_com, _key, com_info, k, url, _stat):
    for curPage in range(1, 100):
        flag, json_data, totalcount = get_baidu_job_json(url, recruitType, projectType, curPage)
        if curPage > total_page:
            break
        if curPage > 5 and _stat['method'] != "cp_full":
            break                          # 非全量模式只抓 5 页

        for item in json_data:
            jobId = item.get("jobId")
            _fullurl = f"https://talent.baidu.com/jobs/detail/{recruitType}/{jobId}"

            # 缓存检查
            _hash = hashlib.md5(_fullurl.encode("utf-8")).hexdigest()
            if os.path.exists(tmp_file) and os.path.exists(tmp_json_file):
                os.utime(tmp_file, ...)    # 更新 mtime
                continue

            # 字段映射:百度 JSON → 内部统一格式
            transform_job_json(item, recruitType, job_type, _key, _fullurl, tmp_file, tmp_json_file)

            # 同时抓取详情页 HTML(供后续大模型抽取使用)
            get_baidu_job_html(_fullurl, tmp_file)

API 路径也要 get_baidu_job_html 下载 HTML 页面。因为后续的 parse_cjob 大模型抽取需要 HTML 格式的职位描述。API JSON 只提供列表字段(标题、地点、部门),职位详情(职责、要求)仍然在 HTML 里。

A.4 字段映射:不管从哪来的,最终都一样

transform_job_json 第 109-151 行:

field_mapping = {
    "announcement_name": "name",
    "publish_time": "publishDate",
    "hd_dept": "bgShortName",
    "hd_loc": "workPlace",
    "hd_job_num": "recruitNum",
    "hd_job_category": "postType"
}
target_json = {}
for target_field, source_field in field_mapping.items():
    target_json[target_field] = item.get(source_field, "")

target_json.update({
    "link": target_url,
    "full_url": target_url,
    "channel": channel,
    "job_type": job_type
})

输出的 detail_*.json 格式和 DOM 路径完全一样。后续的处理流程(SpiderData → parse_cjob)不知道也不关心数据是 DOM 抓的还是 API 拿的。


分支 B:on_response 路径(以兴业银行为例)

B.1 入口

spider_com.py 第 377-378 行:

def on_resp_proc(self, page, _key, com_info, _stat):
    return self._special_proc(page, _key, com_info, _stat, label="on_response", proc=on_response_proc)

auto_on_response/main_proc.py 第 14-21 行:

def on_response_proc(spider_com, page, _key, com_info, k, url, _stat):
    if _key == "com_91000":
        xingye_proc(spider_com, page, _key, com_info, k, url, _stat)

B.2 核心机制:监听浏览器的 HTTP 响应

auto_on_response/xingye_proc.py 第 12-47 行:

def xingye_proc(spider_com, page, _key, com_info, k, url, _stat):
    # 注册响应拦截器
    wrapped_handler = partial(response_handler, spider_com, page, _key, com_info, k, url, _stat, job_type)
    page.on('response', wrapped_handler)

    # 正常打开页面(浏览器会自动发 XHR 请求)
    page.goto(url, timeout=10000)
    time.sleep(10)

    # 翻页:点击"下一页"按钮 → 触发新的 XHR → 拦截器再次捕获
    for i in range(1, _page_count):
        next_page_button = page.get_by_title("下一页")
        if next_page_button and next_page_button.is_enabled():
            aria_disabled = next_page_button.get_attribute('aria-disabled')
            if aria_disabled != 'true':
                next_page_button.click()
            time.sleep(30)

page.on('response', handler) 会拦截页面发出的所有 HTTP 响应——包括 XHR、Fetch、图片、CSS。response_handler 通过 URL 前缀过滤出目标接口:

def response_handler(..., response):
    if response.url.startswith("https://job.cib.com.cn/ersApi/recruitposition/portalPage"):
        _data_json = response.json()
        if _data_json['message'] == '成功':
            for _item in _data_json['data']['list']:
                xingye_json(_item, ...)

这个方案的精妙之处:不需要逆向 API 签名、不需要处理 Token 刷新、不需要模拟请求头——浏览器已经搞定了所有认证。我们只是"偷看"了浏览器自己发出的请求的响应。

B.3 on_response 路径也生成 HTML

xingye_json(第 63-98 行)不仅生成 detail_*.json,还用 API JSON 拼装 HTML

def generate_html(data):
    htmllist = []
    htmllist.append(f"<div> 职位名 {data['positionName']}")
    htmllist.append(f"<div> 工作职责 \n{data['jobDuty']}")
    htmllist.append(f"<div> 任职要求 \n{data['positionRequirment']}")
    return "\n".join(htmllist)

因为后续的大模型抽取需要 HTML。把 API JSON 拼成 HTML,后续流程就不用改。


第四阶段:分页循环

回到 DOM 主干道。get_page_data 处理完第一页后,第 441-469 行的分页逻辑:

_page_count = 0
if "page_func_name" in com_info and _stat['method'] == "cp_full":
    _page_count = 1000        # cp_full 模式:全量翻页
elif "page_count" in com_info and com_info['page_count'] == 'Y':
    _page_count = 3           # 普通模式:只翻 3 页

if _page_count > 1:
    for i in range(2, _page_count):
        _ok = self.pre_page_run(page, com_info, "page_func_name")
        if _ok and i < _page_start:
            continue           # 还没到起始页,跳过
        elif _ok:
            _purl = url + f"&p={i}"
            self.get_page_data(page, _key, com_info, _purl, k)
        else:
            break              # 翻页失败 → 没有下一页了

分页不是"请求第 2 页的 URL"——而是调用 page_func_name 指定的函数来点击"下一页"按钮。因为很多企业网站的翻页不是简单的 ?page=2,可能是 POST 请求、可能是按钮点击触发 JS 加载。

pre_page_run 动态执行 auto_gen_com/gen/gen_xxxxx.py 中的 crawl_page 函数。如果执行失败(按钮不存在或被禁用),循环退出。


第五阶段:从爬取到处理——两条命令的分工

至此,-m cp_full 的爬取阶段完成。磁盘上有了所有公司的 detail_*.html + detail_*.json 文件。

接下来执行第二条命令

python main.py -m cjob -f 2 -d dev

main.py 第 181-184 行:

if args.method == "cjob":
    process_main_announcement(cs, d, _stat, "cjob", "up_api_cjob")

process_main_announcement 第 91-111 行:获取所有公司节点,对每家公司的每个 sch_info 调用:

d.process_announcement_data(_key, _sch_info, _stat, "cjob")

为什么采集和解析分开执行? 因为它们面对的风险完全不同:

  • 采集阶段:网络超时、反爬拦截、页面改版 → 重跑采集即可
  • 解析阶段:模型输出格式错、JSON 解析失败 → 只重跑解析,不用重新抓

分开执行,出问题不会从头再来。而且解析逻辑经常迭代(调 prompt、改字段映射),重跑解析的频率远高于重抓。


第六阶段:SpiderData——数据处理的"路由器"

spider_data.py 第 107-139 行,process_announcement_data 遍历 keydir 下的所有 JSON 文件:

def process_announcement_data(self, key, sch_info, stat, proc_type="ann"):
    keydir = self.spider_sch.get_key_dir(key)
    for file_path in glob.glob(f"{keydir}/*.json"):
        filename = os.path.basename(file_path)
        if filename.startswith("index") or not filename.startswith("detail_"):
            continue

        if self._process_single_file(file_path, filename, key, sch_info, stat,
                                      proc_type, ar_dir, cache_dir):
            break    # 达到处理上限,退出

6.1 _process_single_file:一道文件要过五道关卡

第 141-271 行。每一道关卡都可能让这个文件被跳过:

def _process_single_file(self, file_path, filename, key, sch_info, stat, proc_type, ...):
    # === 关卡 ①:HTML 文件存在性 ===
    hfile = file_path.replace(".json", ".html")
    if proc_type in ["ann", "cjob"] and not os.path.exists(hfile):
        return False                            # 没 HTML → 无法解析

    # === 关卡 ②:10 秒新鲜度 ===
    if check_file_modification_time(file_path):
        return False                            # 可能爬虫还在写这个文件

    # === 关卡 ③:10 天陈旧度 ===
    if check_file_modification_time_old(file_path):
        stat["all_proc_list"].append(tag_info)
        return False                            # 太旧了,之前肯定失败过多次

    # === 关卡 ④:过期标记 ===
    expired_file = ar_file.replace(".json", ".json.expired")
    if os.path.exists(expired_file):
        return False                            # 之前判过死刑

    # === 关卡 ⑤:已处理标记 ===
    model_file = ar_file.replace(".json", ".model.json")
    if os.path.exists(model_file):
        return False                            # 已经处理过了

这五道关卡就是第 5 讲讲的"文件即状态"的完整实践。没有数据库、没有 Redis——文件系统本身就是状态存储。

6.2 cjob 处理分支

第 252-269 行,proc_type == "cjob" 时:

elif proc_type == "cjob":
    ner_logger.info(f"开始处理公司职位文件: {title} - {hfile}")
    ok_cjob, msg = parse_cjob(
        self, model_file, data, sch_info, expired_file, hfile, stat
    )
    if ok_cjob == "ok":
        return self._increment_and_check_stat(stat, "cjob")
    elif ok_cjob == "Err":
        with open(expired_file, "w", encoding="utf-8") as f:
            f.write(msg)                       # 写入失败原因

_increment_and_check_stat 第 273-285 行:每处理一条就计数,达到 DEFAULT_PCOUNT(3 条)就退出。这是一个限流机制——防止单次运行处理太多数据,在开发调试时特别有用。


第七阶段:HTML → 纯文本 → 大模型 → 结构化 JSON

进入 parsegpt/cjob_model.pyparse_cjob(第 22-76 行)。这是数据从"脏 HTML"变成"干净 JSON"的关键一步。

7.1 读取 HTML 并提取正文

def parse_cjob(spider_data, _model_file, _info, com_info, _expired_file, _hfile, _stat):
    with open(_hfile, "r", encoding="utf-8") as f:
        _html = f.read()
        _full_text = get_cjob_html_content(spider_data, com_info, _html)
    _text = clean_text(_full_text)       # 清理多余空白

    if len(_text) < 100:
        return "Err", f"文本长度不足{len(_text)}"   # 内容太短,不浪费时间调模型

get_cjob_html_content(第 155-201 行)是 HTML → 纯文本的关键:

def get_cjob_html_content(spider_data, com_info, htmltext):
    soup = BeautifulSoup(htmltext, 'html.parser')
    fix_html_div(spider_data, soup, com_info, {})    # 移除配置指定的垃圾 div

    class_names = com_info.get("detail_selector")     # 如 "div.content|main"
    if class_names:
        for class_name in class_names.split("|"):
            cc = class_name.split(".")
            div_element = soup.find(cc[0], class_=cc[1]) if len(cc) >= 2 else soup.find(cc[0])
            if div_element:
                return Html2txt().clean_html(str(div_element))

    # 正则回退
    class_name_re = com_info.get("detail_selector_re")
    if class_name_re:
        matched_divs = soup.find_all('div', class_=re.compile(class_name_re))
        if matched_divs:
            return Html2txt().clean_html(str(matched_divs[0]))

    # 全文兜底
    return Html2txt().clean_html(htmltext)

三层提取:精确选择器 → 正则匹配 → 全文Html2txt().clean_html() 是一个自定义的 HTML 转纯文本工具,专门为中文职位内容做了优化(保留换行、处理 <br>、移除 script/style)。

7.2 硬编码字段提取:从 DOM 里直接抠

第 203-240 行,get_hd_element——在调大模型之前,先从 HTML DOM 里直接提取一些确定性字段:

def get_hd_element(htmltext, com_info):
    soup = BeautifulSoup(htmltext, 'html.parser')
    _map = {}

    if com_info.get('detail_hd') == "0001":
        # 百度模式:找 .pos-detail-hd__titBar 里的职位类别
        div_span = soup.find('div', class_='pos-detail-hd__titBar')
        if div_span:
            label_span = soup.find('span', class_='label')
            if label_span:
                _map['hd_job_category'] = label_span.get_text().replace('职位类别:', '')

    if com_info.get('detail_hd') == "0002":
        # 另一家公司的模式:找 .pos-detail-hd__infoBar 里的地点
        ...

    if com_info.get('detail_hd') == "0003":
        # 又一家公司的模式:正则提取发布时间的 yyyy-mm-dd
        match = re.search(r'\d{4}-\d{2}-\d{2}', div_span.get_text())
        if match:
            _map['hd_publish_time'] = match.group()

detail_hd 是一个枚举值,每个值对应一家公司的 DOM 结构模式。这比让大模型从文本里猜"哪段是职位类别"准确得多——DOM 结构是确定的,提取是 100% 准确的。

7.3 调大模型

第 40-44 行:

_t_text = get_template_cjob(_text)      # 系统提示词(定义了 50+ 个字段)
_T_text = get_context_cjob(_text)       # 用户输入(待抽取的文本)
(_ok_flag, json_str, tokens) = new_call_gpt(_t_text, _T_text, True)

get_template_cjob 返回一个约 400 行的提示模板(template.py 第 212-425 行),定义了职位抽取的全部字段规范。这个模板的体量说明了大模型抽取的复杂性——不是"把文本丢给 GPT 让它提取",而是需要极其详细的字段定义、提取规则、标准化规范。

new_call_gpt 使用的是 DeepSeek 的 Responses API,支持 previous_response_id 上下文缓存——相同前缀的请求可以复用缓存的 KV cache,大幅降低 token 开销。

7.4 字段回填:规则修正模型输出

第 79-152 行,set_other_info——大模型输出后,用规则修正和补全:

def set_other_info(com_info, _info, json_data, _text, _a_map):
    json_data['FileId'] = getMD5Str(_text)
    json_data['JobLink'] = _info['full_url']
    json_data['JobTitle'] = fix_job_name(_info['announcement_name'])
    json_data['ComName'] = com_info['com_name']          # 从配置拿,不靠模型
    json_data['ComShortName'] = com_info['com_webname']
    json_data['DocType'] = _info['job_type']

    # 从列表数据回填(列表页已经有的信息,不靠模型猜)
    if 'hd_loc' in _info and len(_info['hd_loc']) > 1:
        json_data['WorkPlace'] = _info['hd_loc']
    if 'publish_time' in _info and len(_info['publish_time']) > 1:
        json_data['PublishTime'] = _info['publish_time']

    # 从 DOM 硬编码提取回填
    if 'hd_loc' in _a_map and len(_a_map['hd_loc']) > 0:
        json_data['WorkPlace'] = _a_map['hd_loc']

    # 没有发布日期 → 用当前日期
    if not 'PublishTime' in json_data or json_data['PublishTime'] == '':
        json_data['PublishTime'] = get_current_data()

字段来源的优先级

DOM 硬编码提取 > 列表数据回填 > 大模型抽取 > 规则兜底
    100%准确      99%准确       语义理解      补缺

7.5 质量校验

第 52-54 行:

if len(json_data['JobDescribe']) + len(json_data['Jobreq']) < 30:
    return "Err", f"职位信息的描述太少:{_hfile}\n{json_str}"

30 个字符的判断——如果职位描述和要求加起来不足 30 个字符,说明要么页面是空的,要么大模型没提取到实质内容。直接标记为 Err,写 expired_file

7.6 最终落盘

第 72-76 行:

with open(_model_file, 'w', encoding='utf-8') as fw:
    json.dump(_ann_dict, fw, ensure_ascii=False, indent=4)

model.json 的结构:

{
  "cjob": {
    "JobTitle": "Java开发工程师",
    "ComName": "某某公司",
    "WorkPlace": "北京",
    "Salary": "20k-40k",
    "JobDescribe": "1. 负责后端服务开发...",
    "Jobreq": "1. 本科及以上学历...",
    "Degree": ["本科及以上"],
    "TypeAndLevel": {"Level": "L2", "Codes": ["1302"]},
    "Skills": ["Java", "Spring Boot", "MySQL"],
    ...
  },
  "other": {
    "announcement_name": "Java开发工程师",
    "full_url": "https://example.com/job/12345",
    "channel": "com_00003",
    "process_time": "2026-05-08T14:30:00",
    ...
  }
}

完整生命线回顾

让我们把一条"Java开发工程师"职位从生到死的完整路径画出来:

python main.py -m cp_full -f 2 -d dev
│
├─ ① main.py:315 → args.method="cp_full", args.file="2"
├─ ① main.py:321 → cs = SpiderCom("2")
│   └─ spider_com.py:42-64
│       setting_default.ini → setting_template.ini → setting_com_2.ini
│
├─ ① main.py:337run_periodically(s, cs, d)
├─ ① main.py:175clawler_main(cs, _stat)
│   └─ main.py:41-71sync_playwright() → browser → page → get_nodes() → get_progress()
│
├─ ② main.py:67 → cs.run(page, "com_00003", com_info, _stat)
│   └─ spider_com.py:406
│       data_proc_type == ""DOM 路径
│
├─ ③ spider_com.py:435open_with_url(page, url)
│   └─ page.goto(列表URL) → wait load → wait networkidle
│
├─ ③ spider_com.py:482get_selector_text(主选择器 → 备选列表 → 正则)
│   └─ 找到 ".job-list-container"
│
├─ ③ spider_com.py:489_auto_scroll_to_bottom (判停: 高度不变)
├─ ③ spider_com.py:490_click_load_more (element/function 两种模式)
│
├─ ③ spider_com.py:492-505 → tableObj.inner_html() → index_*.html
├─ ③ spider_com.py:513call_func("auto_gen_com.gen.gen_00010", html)
│   └─ 动态导入 → extract_table_from_html() → index_*.json
│       结果为空? → gen_func_bygpt() 自动生成新解析函数
│
├─ ③ spider_com.py:525 → 循环 index_*.json 每条职位
│   └─ get_page_detail_data(page, ..., _item)
│       ├─ 处理链接:  → 清空; 为空 → 点击获取; 相对 → 补全
│       ├─ 检查缓存: detail_*.html + detail_*.json 存在? → 更新 mtime, return
│       ├─ 拼接 URL: get_full_url(domain, link)
│       ├─ 前端路由保护: if "#" in url → 跳过 get_final_url
│       ├─ 打开详情: get_page_detail_content → browser.new_page() 新 tab
│       │   ├─ page.goto(detail_url) → 等 networkidle
│       │   ├─ 关弹窗: "Accept"/"同意"/"Continue"click(timeout=1500)
│       │   └─ 正文提取: detail_selector → 正则 → 全文
│       ├─ 落盘: detail_*.html + detail_*.json
│       └─ monitor.log_crawl(成功/失败)
│
├─ ④ spider_com.py:452-469 → 分页循环
│   └─ pre_page_run → page_func_name → 点击下一页 → get_page_data(新URL)
│
├─ ③ spider_com.py:473write_process_file("com_00003")
│
└─ 采集阶段结束。磁盘上:
    data/tmp/com_00003/
    ├── index_abc.html           ← 列表页 HTML
    ├── index_abc.json           ← 解析出的 50 条职位列表
    ├── detail_def.html + .json  ← 第 1 条职位
    ├── detail_ghi.html + .json  ← 第 2 条职位
    └── ...                      ← 共 50 对文件

═══════════════════════════════════════════
第二条命令:python main.py -m cjob -f 2 -d dev
═══════════════════════════════════════════

├─ ⑤ main.py:181process_main_announcement(cs, d, _stat, "cjob")
│   └─ SpiderData.process_announcement_data("com_00003", sch_info, _stat, "cjob")
│
├─ ⑥ spider_data.py:126 → 遍历 glob("data/tmp/com_00003/*.json")
│   对每个 detail_*.json:
│
├─ ⑥ _process_single_file:
│   ├─ 关卡①: HTML 文件存在?
│   ├─ 关卡②: 10 秒内生成? → skip (防并发)
│   ├─ 关卡③: 10 天前? → skip (太旧)
│   ├─ 关卡④: .expired 存在? → skip (已判死刑)
│   ├─ 关卡⑤: .model.json 存在? → skip (已处理)
│   └─ proc_type == "cjob"parse_cjob(...)
│
├─ ⑦ parsegpt/cjob_model.py:28-32:
│   ├─ 读 detail_*.html
│   ├─ get_cjob_html_content: 提取正文块 (选择器 → 正则 → 全文)
│   ├─ clean_text: 清理多余空白
│   └─ 文本 < 100 字符? → Err
│
├─ ⑦ get_hd_element: DOM 硬编码提取 (detail_hd 枚举驱动)
│   └─ 职位类别、地点、发布日期
│
├─ ⑦ template.py:436get_template_cjob(_text)
│   └─ 拼装 400 行提示词(50+ 字段定义、标准化规则、示例)
│
├─ ⑦ new_call_gpt(系统提示词, 待抽取文本) → 大模型返回 JSON
│
├─ ⑦ set_other_info: 字段回填
│   ├─ ComName ← 配置里的 com_name
│   ├─ WorkPlace ← 列表数据 hd_loc / DOM 提取 hd_loc
│   ├─ PublishTime ← 列表数据 / DOM 提取 / 当前日期(兜底)
│   └─ DocTypejob_type (shezhao/xiaozhao/shixi)
│
├─ ⑦ 质量校验:
│   ├─ JobDescribe + Jobreq < 30 字符? → Err
│   └─ Degree 字段规范化
│
└─ ⑦ 落盘: detail_*.model.json
    {
      "cjob": { /* 50+ 结构化字段 */ },
      "other": { /* 元信息 */ }
    }

关键设计决策汇总

你在代码里看到的 背后解决的问题 如果不用这个设计
configparser 三层 read 全局默认 + 模板 + 分片特定,灵活覆盖 每个分片都要写完整配置,几百个 INI 文件没法维护
supplement_node_infoif not _key in 公司自身配置优先于模板 模板一改,所有公司的个性化配置都被覆盖
get_selector_text 三层回退 站点改版时 CSS 类名变化 主选择器失效 → 100% 失败,只能人工修
列表 HTML 先落盘再解析 数据安全和可重跑 解析函数写错了,原始数据也丢了
.url 缓存文件 避免重复点击获取链接 每次重跑都要操作 DOM 点击,慢且不稳定
if "#" not in _fullurl 前端路由 URL 不能被 HTTP 重定向破坏 所有 SPA 站点的详情链接全部失效
新开 tab 访问详情页 隔离列表页和详情页的 JS 上下文 返回列表页时触发重新加载,丢滚动位置
正文三层提取(选择器→正则→全文) 应对详情页 HTML 结构多样性 经常因为找不到正文容器而失败
data_proc_type 一个字段切三通道 不同站点用最优采集方式,但对外接口统一 要么全用 DOM(慢),要么全用 API(很多站没有)
字段来源优先级(DOM > 列表 > 模型 > 兜底) 能用精确规则的不用概率模型 模型幻觉引入错误数据
check_file_modification_time 10 秒过滤 防解析进程和采集进程并发冲突 读到写了一半的文件,JSON 解析报错
采集和解析分两条命令执行 隔离不同失败模式,支持独立重跑 采集时解析失败→重跑要从头抓,浪费时间
file_path 存在即状态 零依赖的状态追踪 需要额外维护数据库或状态文件
解析失败自动 gen_func_bygpt 列表页改版后零人工介入尝试修复 每次改版都要人工分析 HTML、写新解析函数
提示模板 400 行 大模型需要极其详细的字段定义才能输出一致格式 每次输出字段名和格式都不一样,无法入库

本讲核心认知

追踪一条数据的生命线,你会发现这个系统没有"架构图上的漂亮模块",只有踩过几百个坑之后长出来的防御性代码。

get_selector_text 的三层回退不是设计出来的,是一个又一个站点改版后加上的。

if "#" not in _fullurl 这一行不是"最佳实践",是某个 SPA 站点全部失效后打上的补丁。

open_with_urlstatus in [200, 412] 的 412,背后是兰州大学那个奇怪的服务器配置。

工程代码的价值不在"优雅",在"每一种会出错的情况都有对应的一行代码在兜底"。


组内导航:⬅️ 从能跑到能稳定跑很久——爬虫的工程化跃迁 | 🏠 00-爬虫课程六讲