当前位置:首页>python>Python 爬取 B 站 UP 主所有动态,真正麻烦的是第二页

Python 爬取 B 站 UP 主所有动态,真正麻烦的是第二页

  • 2026-09-17 15:28:49
Python 爬取 B 站 UP 主所有动态,真正麻烦的是第二页
第一页能拿到数据,第二页却开始重复。

这种代码我第一眼一般不看请求头,先看翻页参数。B 站动态列表不是常见的 page=1、page=2,而是游标翻页:第一次请求不传 offset,下一次请求使用上一次响应里的 offset。什么时候 has_more 变成 false,什么时候才算拉完。

当前网页端使用的动态接口大致是:

GET /x/polymer/web-dynamic/v1/feed/space

主要参数只有两个:

host_mid:UP 主 UID
offset:上一页返回的游标

别自己计算 offset,也别拿动态 ID 加一。这个值由服务端返回,看着像一串没规律的数字,原样带回去就行。

下面这份代码没有上代理池、UA 池那套东西。抓一个 UP 主的公开动态,控制请求频率、处理翻页、保留原始数据,够用了。

from __future__ import annotations

import json
import os
import random
import time
from datetime import datetime
from pathlib import Path
from typing import Any

import requests


DYNAMIC_API = (
"https://api.bilibili.com/"
"x/polymer/web-dynamic/v1/feed/space"
)


classUpDynamicCollector:
def__init__(self, uid: str, cookie: str = "") -> None:
        self.uid = uid
        self.client = requests.Session()

        self.client.headers.update({
"User-Agent": (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 Chrome/131.0 Safari/537.36"
            ),
"Referer": f"https://space.bilibili.com/{uid}/dynamic",
"Accept": "application/json, text/plain, */*",
        })

if cookie:
            self.client.headers["Cookie"] = cookie

defrequest_page(self, offset: str | None) -> dict[str, Any]:
        params = {"host_mid": self.uid}

if offset:
            params["offset"] = offset

        last_error: Exception | None = None

for attempt in range(3):
try:
                response = self.client.get(
                    DYNAMIC_API,
                    params=params,
                    timeout=(5, 20),
                )
                response.raise_for_status()

                payload = response.json()
if payload.get("code") != 0:
raise RuntimeError(
f"接口返回异常:{payload.get('code')} "
f"{payload.get('message')}"
                    )

return payload["data"]

except (requests.RequestException, ValueError, RuntimeError) as exc:
                last_error = exc
                time.sleep(2 + attempt * 3)

raise RuntimeError(f"连续请求失败:{last_error}")

    @staticmethod
defparse_item(item: dict[str, Any]) -> dict[str, Any]:
        modules = item.get("modules") or {}
        author = modules.get("module_author") or {}
        dynamic = modules.get("module_dynamic") or {}
        desc = dynamic.get("desc") or {}
        stat = modules.get("module_stat") or {}

        publish_time = author.get("pub_ts")
if publish_time:
            publish_time = datetime.fromtimestamp(
                publish_time
            ).astimezone().isoformat()

return {
"dynamic_id": item.get("id_str"),
"dynamic_type": item.get("type"),
"author": author.get("name"),
"publish_time": publish_time,
"text": desc.get("text", ""),
"forward_count": (
                stat.get("forward", {}).get("count", 0)
            ),
"comment_count": (
                stat.get("comment", {}).get("count", 0)
            ),
"like_count": (
                stat.get("like", {}).get("count", 0)
            ),
# 图片、视频、专栏、投票等结构并不统一。
# 原始对象必须留下,后面改解析逻辑时不用重新爬。
"raw": item,
        }

defcollect(self, output_file: str) -> int:
        output = Path(output_file)
        output.parent.mkdir(parents=True, exist_ok=True)

        offset: str | None = None
        seen_ids: set[str] = set()
        total = 0

with output.open("w", encoding="utf-8") as file:
whileTrue:
                page = self.request_page(offset)
                items = page.get("items") or []

ifnot items:
break

for item in items:
                    dynamic_id = str(item.get("id_str", ""))

# 置顶动态有时会在后续页面再次出现。
ifnot dynamic_id or dynamic_id in seen_ids:
continue

                    seen_ids.add(dynamic_id)
                    record = self.parse_item(item)

                    file.write(
                        json.dumps(record, ensure_ascii=False) + "\n"
                    )
                    total += 1

                print(
f"本页 {len(items)} 条,"
f"累计写入 {total} 条"
                )

ifnot page.get("has_more"):
break

                next_offset = page.get("offset")
ifnot next_offset or next_offset == offset:
                    print("offset 没有变化,停止翻页")
break

                offset = next_offset
                time.sleep(random.uniform(1.8, 3.5))

return total


if __name__ == "__main__":
    up_uid = "替换成UP主UID"

# Cookie 不要直接写进代码,更不要提交到 Git 仓库。
    bili_cookie = os.getenv("BILI_COOKIE", "")

    collector = UpDynamicCollector(
        uid=up_uid,
        cookie=bili_cookie,
    )

    count = collector.collect(
        output_file=f"data/{up_uid}_dynamic.jsonl"
    )
    print(f"抓取结束,共保存 {count} 条动态")

运行前安装依赖:

pip install requests

登录态 Cookie 放到环境变量里:

export BILI_COOKIE='这里放你自己的Cookie'
python collect_dynamic.py

Windows PowerShell 可以这样写:

$env:BILI_COOKIE="这里放你自己的Cookie"
python collect_dynamic.py

我这里没有直接保存成 CSV,而是用了 JSONL,也就是一行一条 JSON。

原因很实际。B 站动态类型不少,纯文字、图片、视频投稿、专栏、转发动态的数据结构都不完全一样。现在为了表格好看,强行把字段拍平,过两天想补图片地址或者转发原文,还得重新爬一遍。

所以抓取阶段只做两件事:提取常用字段,保留完整的 raw。

还有一个容易漏的地方是置顶动态。它可能在第一页出现,后续翻页时又碰到一次。代码里用 dynamic_id 做了去重,不然最后统计数量会比实际动态数多。

接口出现 -352、-412 或 HTTP 429 时,也别马上套代理继续冲。先停下来检查 Cookie、请求频率和接口是否已经调整。这个接口来自网页端调用,并不是一个承诺长期兼容的稳定开放接口;B 站同时提供正式开放平台,长期项目更适合先确认开放平台是否有对应能力。

爬虫最怕的不是某次请求失败,而是失败以后悄悄少了几十页,程序最后还打印一句“抓取成功”。

所以超时、重试、游标不变检测、原始数据留存,这几个地方不能省。请求慢一点没事,数据缺一截,后面补起来才麻烦。

最新文章

随机文章