本文へ移動
cccskills
無料GitHub で公開

ra-公众号提取

提取微信公众号文章全文。

Use when the user provides a mp.weixin.qq.com URL and asks for 提取文章, 抓取公众号, 读一下这篇, or when another skill (ra-洗稿, ra-选题, ra-video-wash-pipeline) needs to read a WeChat article before processing. Uses MicroMessenger UA spoofing to bypass WeChat's first-layer access control. Stdlib only, no API key required.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md4.5 KB
  • scripts/fetch_wechat.py6.6 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

ra-公众号提取

Goal

Extract the full text of a WeChat public account article into clean Markdown, ready for downstream skills.

This skill is intentionally narrow:

  • fetch the article HTML with a spoofed WeChat UA
  • parse title, author, publish time, and body text
  • write a Markdown file to 01-内容生产/视频工作台/.internal/公众号/
  • print the complete text in the conversation

Do not summarize, rewrite, translate, title, or start production. Hand the extracted text to the calling skill or user.

Locate The Skill

WX_HOME="$(
  for d in "$(pwd)/.codex/skills/ra-公众号提取" \
           "$(pwd)/.agents/skills/ra-公众号提取" \
           "$(pwd)/.claude/skills/ra-公众号提取" \
           "$HOME/.codex/skills/ra-公众号提取" \
           "$HOME/.agents/skills/ra-公众号提取" \
           "$HOME/.claude/skills/ra-公众号提取"; do
    [ -f "$d/SKILL.md" ] && echo "$d" && break
  done
)"
export WX_HOME

Use "$WX_HOME/scripts/fetch_wechat.py" for every operation.

How It Works

WeChat blocks external access to public account articles through several layers:

  1. User-Agent detection — checks for MicroMessenger keyword; rejects normal browser UAs
  2. Referer check — expects requests from mp.weixin.qq.com
  3. JS lazy-loading — images use data-src instead of src
  4. Rate limiting — frequent requests trigger CAPTCHA

This skill bypasses layer 1 and 2 by sending a request with a genuine WeChat iOS WebView User-Agent and the correct Referer header. This is sufficient to receive the full article HTML for text extraction. No login, cookie, or API key is needed.

Limitations:

  • Images are extracted as URLs only (from data-src), not downloaded
  • Some articles with heavy JS rendering may return partial content
  • Rapid successive calls from the same IP may trigger CAPTCHA (wait and retry)
  • Does not work on articles that require WeChat login or payment

Workflow

  1. Health check (optional):

    python3 "$WX_HOME/scripts/fetch_wechat.py" --doctor
    
  2. Extract article:

    python3 "$WX_HOME/scripts/fetch_wechat.py" "<mp.weixin.qq.com URL>" \
      --output-dir "01-内容生产/视频工作台/.internal/公众号"
    
  3. The script outputs a JSON result with title, author, publish_time, char_count, and output_path. Read the output file and present the full text to the user or pass it to the next skill.

  4. If the script exits with code 1 (fetch/CAPTCHA failure), report the error and suggest:

    • Wait a few minutes and retry
    • Ask the user to paste the article text directly
  5. If the script exits with code 2 (parse failure), save the raw HTML with --raw for debugging.

Storage Rules

  • Extracted articles land in 01-内容生产/视频工作台/.internal/公众号/ by default (source privacy zone).
  • Source URLs and source titles must not appear in any handoff file, finished script, or public document.
  • When called by ra-洗稿 or ra-选题, the source link is recorded only in the topic card's 来源: field.

Integration With Other Skills

This skill is a source reader, not a content processor. Typical call chains:

User saysCall chain
这篇公众号洗稿 / 基于这篇公众号制作视频ra-公众号提取 → ra-洗稿
这篇文章存个选题ra-公众号提取 → ra-选题
读一下这篇公众号ra-公众号提取 (standalone)
这篇公众号做成图文ra-公众号提取 → ra-洗稿 (图文形态)

When another skill needs to read a mp.weixin.qq.com URL, call this skill first to obtain the text, then pass the extracted Markdown to the downstream skill.

Script Reference

fetch_wechat.py <url> [options]

Arguments:
  url                    mp.weixin.qq.com article URL

Options:
  --output-dir <dir>     Output directory (default: current directory)
  --raw                  Also save the raw HTML alongside the Markdown
  --doctor               Check dependencies and exit

Exit codes: 0 = success, 1 = fetch failed, 2 = parse failed.

No external dependencies — uses Python stdlib only (urllib, re, html, json).

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

口播视频转录和口误识别。生成审查稿和删除任务清单。触发词:剪口播、处理视频、识别口误

日本語の概要は準備中です。原文の説明を表示しています。

Pluviobyte/rnskill1,6432026年9月21日 更新

将法天象地、同源巨大法相或角色力量显现的视觉参考制作成连续关键帧,再使用用户选择或可用的图生视频工具分段生成、检查和合成动作视频,不绑定成片平台。适用于本体与法相同框、领域展开、凝实与同步出招;不负责整部小说改编或普通图片轮播。

日本語の概要は準備中です。原文の説明を表示しています。

Pluviobyte/rnskill1,6432026年9月21日 更新

口播基础素材包生成。转录口播、识别口误、生成审核页;用户确认后剪出新视频,Agent 再基于剪后视频重新转写、AI 校对字幕,输出后续口播成片可用的 source_cut.mp4 和 subtitles.srt。触发词:剪口播、处理口播素材、准备口播素材、识别口误、基础素材包

日本語の概要は準備中です。原文の説明を表示しています。

Pluviobyte/rnskill1,6432026年9月21日 更新

口播视频成片 Skill。把文章/口播稿/SRT、剪后视频和 HTML/图片素材串成分镜稿、时间线预览和最终 MP4;成片比例和动画风格从用户配置读取,动画默认使用小黑风格。触发词:口播成片、做分镜稿、时间线预览、合成口播视频、导出竖屏MP4

日本語の概要は準備中です。原文の説明を表示しています。

Pluviobyte/rnskill1,6432026年9月21日 更新

自进化 skills。记录用户反馈,更新方法论和规则。触发词:更新规则、记录反馈、改进skill

日本語の概要は準備中です。原文の説明を表示しています。

Pluviobyte/rnskill1,6432026年9月21日 更新

dbs

無料

dontbesilent 商业工具箱主入口。双模式:任务前路由(你的问题该用哪个 skill)+ 任务后导航(刚做完诊断,下一步该干什么)。 触发方式:/dbs、/商业、「帮我看看」、「下一步怎么走」 Main entry point for dontbesilent business toolkit. Dual mode: pre-task routing + post-task navigation. Trigger: /dbs, "help me with my business", "what's next"

日本語の概要は準備中です。原文の説明を表示しています。

Pluviobyte/rnskill1,6432026年9月21日 更新

Pluviobyte のスキルをすべて見る

このスキルの問題を報告する