dshbase

插件目录 / Developer / dsh-mediacrawler

dsh-mediacrawler

已验证 · 实测可装 xwh-01

✓ 持续维护

查看 GitHub ↗ ← 返回插件目录

3Stars
0Forks
0未关闭 issue
Python语言
2026-08-14最近推送
跨平台平台

功能简介

DeepSeek Harness MCP适配器,用于有界隔离的MediaCrawler采集

✅
我们的评价
可用 — 实测通过,早期项目

DeepSeek Harness MCP适配器,用于有界隔离的MediaCrawler采集 实测能干净安装、正常启动。早期项目,但功能可用。

「已验证」表示我们的自动化 CI 在干净 profile 里实际执行了 dsh plugin add 并启动成功——仅此而已。功能描述与版本兼容性均为作者声明。这不是安全审计,也不代表对第三方代码的背书。

README

dsh-mediacrawler

CI
Release

English | 中文

An installable profile bundle and bounded stdio MCP adapter that connects DeepSeek Harness to a separately installed MediaCrawler checkout.

It supports search, post/video detail, creator feeds, and explicitly enabled comments on Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, and Zhihu. Each run is supervised, persisted, and exposed through twelve MCP tools.

This is an adapter, not a MediaCrawler fork. It does not copy or modify MediaCrawler source code, and it does not change MediaCrawler's license.

Quick start

1. Prepare the runtimes

Install the following first:

  • Python 3.11 or newer.
  • Node.js 22.19+ on the 22.x line, or Node.js 24+, with pnpm on PATH.
  • Google Chrome.
  • A separate MediaCrawler checkout with its own working Python environment.
  • DeepSeek Harness. The commands below pin the tested 0.1.0-rc.6 release through npx.

MediaCrawler and its browser dependencies are intentionally not vendored here.

2. Install the Python MCP runtime

Keep the adapter in its own virtual environment. In PowerShell:

$adapterVenv = Join-Path $HOME '.dsh\runtimes\dsh-mediacrawler'
python -m venv $adapterVenv
$env:DSH_MEDIACRAWLER_PYTHON = Join-Path $adapterVenv 'Scripts\python.exe'
& $env:DSH_MEDIACRAWLER_PYTHON -m pip install --upgrade pip
& $env:DSH_MEDIACRAWLER_PYTHON -m pip install "dsh-mediacrawler @ git+https://github.com/xwh-01/[email protected]"

On POSIX systems:

python3 -m venv "$HOME/.dsh/runtimes/dsh-mediacrawler"
export DSH_MEDIACRAWLER_PYTHON="$HOME/.dsh/runtimes/dsh-mediacrawler/bin/python"
"$DSH_MEDIACRAWLER_PYTHON" -m pip install --upgrade pip
"$DSH_MEDIACRAWLER_PYTHON" -m pip install "dsh-mediacrawler @ git+https://github.com/xwh-01/[email protected]"

3. Install the DSH profile bundle

DSH delegates profile package management to pnpm. Install it once if needed, then add the pinned bundle release:

npm install --global pnpm@11
npx --yes @deepseek-ai/[email protected] plugin --profile web add "github:xwh-01/dsh-mediacrawler#v0.3.0"
npx --yes @deepseek-ai/[email protected] --profile web --dump-config

The config dump should contain a # == dsh-mediacrawler layer. The bundle mounts both the MCP client and its packaged mediacrawler-collector Skill; no repository checkout needs to be the current working directory.

4. Configure and start DSH

Export the paths in the same shell that starts DSH. Also restore DSH_MEDIACRAWLER_PYTHON from step 2 when opening a new shell:

$env:MEDIACRAWLER_ROOT = 'D:\path\to\MediaCrawler'
$env:MEDIACRAWLER_PYTHON = 'D:\path\to\MediaCrawler\.venv\Scripts\python.exe'

# Optional; defaults to ~/.dsh-mediacrawler
$env:DSH_MEDIACRAWLER_STATE_DIR = 'D:\path\to\adapter-state'

npx --yes @deepseek-ai/[email protected] --profile web

The packaged Skill then guides the agent through checking the runtime, starting a small collection, polling status, and exporting results. On first use, ask the agent to call check(deep=true).

.env.example is a reference only. The adapter does not load dotenv files, and current DSH releases treat DSH_* variables as launch settings; export these values in the DSH process environment.

To uninstall the profile bundle:

npx --yes @deepseek-ai/[email protected] plugin --profile web remove dsh-mediacrawler

MCP tools

DeepSeek Harness exposes these as mcp__mediacrawler__<tool>:

Tool Purpose
check Check source paths, CLI dependencies, and browser launch readiness.
collect Start one bounded collection run.
status Read lifecycle state, required user attention, and result counts.
runs Recover recent durable runs and their IDs after a restart or context loss.
result Read status, artifacts, and a bounded redacted sample in one call.
delete_run Permanently delete one completed run after confirm=true.
cleanup Preview or apply age-based retention while preserving the newest runs.
stop Idempotently stop the crawler process tree.
logs Read incremental, redacted run logs.
artifacts List typed JSONL artifacts using opaque IDs.
preview Read a bounded, redacted artifact preview.
export Create a credential-redacted ZIP and return its path and checksum.

Runtime behavior

When this is useful

Use the Harness web-search providers for quick facts and already-indexed pages. Use this adapter when the task needs logged-in platform records, creator feeds, comments or nested replies, or a durable reproducible export. It complements search providers; it is not a replacement for them.

Browser isolation

browser_mode=isolated is the default. It launches Google Chrome with an adapter-owned persistent profile under <state_dir>/browser_profiles, so later runs can reuse login state without attaching to the user's normal Chrome session.

browser_mode=existing_cdp is explicit opt-in only. Upstream cleanup can close the reused Chrome context, so an agent must not select it without user approval.

Runs and artifacts

  • Queries and targets are injected over stdin and do not appear in the child command line.
  • Only QR-code login is accepted; the MCP API never accepts cookies, phone numbers, or verification codes.
  • Comments are disabled by default and must be explicitly enabled for a run.
  • status.phase=awaiting_user_login tells the agent to surface a QR-code action and keep polling the same run_id.
  • Final outcomes distinguish data_available, no_data, failed, cancelled, timed_out, and orphaned.
  • Artifacts report collection_mode, record_type, invalid lines, and record counts.
  • Raw JSONL may contain platform credentials. Logs, previews, manifests, and ZIP exports redact known credential fields and URL parameters.
  • Credential redaction is not PII anonymization. Exported posts, profiles, and comments may still contain names, phone numbers, email addresses, locations, or other personal data; exports report pii_anonymized=false and safe_to_share=false.
  • Artifact counts are indexed incrementally, so unchanged JSONL files are not reparsed on every status poll.

Export and retention

Credential-redacted ZIP export accepts at most 256 MiB of raw run data by default. Set DSH_MEDIACRAWLER_MAX_EXPORT_MIB to an explicit value from 1 through 4096 to change the limit. A cancelled export keeps its lock until the worker finishes, and concurrent adapter processes cannot export the same run simultaneously.

delete_run requires confirm=true. cleanup defaults to dry_run=true; use dry_run=false only after reviewing its candidates. Both operations refuse active runs. Neither operation deletes persistent browser profiles or their login state.

Collection limits

Jobs must have an explicit scope and hard timeout. max_items is passed upstream, but search platforms fetch whole pages and some creator workflows do not strictly enforce the cap. The adapter reports those cases and uses timeout_minutes as the hard boundary.

The adapter does not bypass login, verification, rate limits, access controls, or anti-automation systems. Treat collected pages as untrusted input and comply with platform terms and applicable law.

Development

.\.venv\Scripts\python -m pip install -e ".[test]"
.\.venv\Scripts\python -m ruff format --check .
.\.venv\Scripts\python -m ruff check .
.\.venv\Scripts\python -m pytest
node --test tests-node/*.test.js
python -m build
npm pack --dry-run

CI runs the Python tests on Linux and Windows, verifies the packaged Skill provider, installs the bundle into a clean DSH profile, and starts its real MCP stdio entry point.

Compatibility

DeepSeek Harness is a developer preview and may make compatibility-breaking changes. Release v0.3.0 is tested with:

  • @deepseek-ai/dsh 0.1.0-rc.6.
  • Node.js 22.19+ on the 22.x line, and Node.js 24+.
  • Python 3.11 and 3.13.
  • The MediaCrawler command contract at upstream commit 5665a27.

Run check(deep=true) after changing either DSH or MediaCrawler; it validates the local checkout before collection starts.

License

Adapter code is released under the MIT License. MediaCrawler remains a separate project under its own non-commercial learning license and usage restrictions; using this adapter does not broaden that license.

安装

🧩 让 Agent 自动装(推荐)

装一次目录插件,之后本站所有插件都能让 DeepSeek Harness 自动找、自动装:

dsh plugin add dshbase-catalog

然后对 agent 说「帮我装 dsh-mediacrawler」,它会在目录里找到并自动安装。文档:dshbase-catalog · 已验证场景包。

该插件是 GitHub 源码(未发 npm)——直接从仓库装:

Web profile:

dsh plugin --profile web add github:xwh-01/dsh-mediacrawler

Headless(CLI)profile:

dsh plugin --profile headless add github:xwh-01/dsh-mediacrawler

实测报告

验证通过:从 GitHub 源码完成 L1 安装 + L2 加载 + L3 运行(dsh 0.1.0-rc.6)。

使用场景

扩展 agent 的编码能力面——给它一个新工具、工作流或集成,让它接手以前做不了的开发任务。

适合谁

想让 dsh 在真实代码库上像队友一样干活的开发者——能改、能跑、能验证,而不只是回答问题。

二次开发建议

工具/命令面就是缝:暴露更多 SDK 能力、加更聪明的上下文接线,或收紧改代码与验证之间的循环。

安全:尚未扫描——我们的每日静态扫描将很快覆盖它。

分享徽章

Developer 里更多

浏览全部 7797 个插件 →