| name | topic-to-digital-human-video |
|---|---|
| description | Turn a concrete topic into a researched, fact-checked, captioned Chinese digital-human talking-head video using the local DuckDuckGo fast-research API, Qwen3-TTS voice workflow, HeyGen presenter bookends, real source screenshots, and Remotion. Use when the user supplies a topic and wants a finished 口播视频/数字人口播成片, or wants to resume, revise, or quality-check any stage of that topic-to-video pipeline. |
Topic to Digital-Human Video
Purpose
Convert one concrete topic into a publishable 16:9 Chinese talking-head explainer. Treat research, script, voice, presenter footage, source visuals, captions, cover, and final QA as one product workflow rather than unrelated tools.
Default to the installed local stack and the user's authorized voice and portrait. Discover paths from TOPIC_VIDEO_WORKSPACE and the optional component-specific environment variables documented in references/local-stack.md. Do not ask the user to repeat configuration that can be discovered locally.
Prerequisite Gate
Check the local stack before accepting the topic as ready for production. GPT Researcher is a required local service, not an optional enhancement.
Run scripts/check_stack.py first and handle its prerequisite.state exactly:
install_required: pause and tell the user that the customized GPT Researcher service must be installed before this Skill can research or verify a topic. Offer to install or restore it; do not silently clone or configure software.service_start_required: tell the user the service is installed but not running, start it with the command returned by the check when authorized by the request, and wait until the route is healthy.custom_api_required: explain that the running service lacks this workflow's customized/api/script/packageendpoint and must be updated before continuing.dependencies_missingorstack_restore_required: name the missing components and pause for installation or restoration.ready: continue with the topic.
Use this user-facing message when installation is required: 开始前需要先安装本地 GPT Researcher 调研服务(包含 DuckDuckGo 快速调研和 /api/script/package 接口)。安装完成并启动后,我才能自动完成事实核验、脚本和真实来源截图。
Do not claim that installing the unmodified upstream repository alone provides /api/script/package; this workspace uses a customized local API.
Required References
Read both files before starting production:
references/local-stack.mdfor paths, API payloads, commands, and artifact locations.references/quality-gates.mdfor evidence rules, paid-action approval, revision routing, and final QA.
When the task enters voice, HeyGen, or final caption rendering, also follow the installed digital-human-video-production and local-video-captions skills if available. This orchestrator adds the topic-research and complete-product layer; it does not relax their safety or quality requirements.
Default Input Contract
Only topic is required. Unless the user overrides them, use:
- audience:
普通中文用户 - duration: 60 seconds, with 45–90 seconds acceptable after evidence-driven editing
- script variants: 3
- research sources: up to 15
- real source screenshots: 4–6
- format: 1280×720, 16:9, H.264 video plus AAC audio
- structure: cover → digital-human opening → real source visuals → digital-human ending → 3-second subscribe/like card
Accept optional hypotheses, audience, target duration, platform, tone, and budget. Treat hypotheses as claims to test, never as facts to preserve.
Workflow
1. Inspect and initialize
Complete the prerequisite gate above. Restart only a stopped local service; do not reinstall or reconfigure healthy components.
Initialize exactly one topic directory with scripts/init_production.py. Store everything beneath:
$TOPIC_VIDEO_WORKSPACE/productions/<topic>/
Never place generated topic files in the workspace root. Reuse an existing topic directory when its production.json matches.
2. Research by API, not by web form
Call POST $TOPIC_VIDEO_RESEARCH_API/api/script/package directly, defaulting the API base to http://127.0.0.1:8000. This fast path uses DuckDuckGo search, parallel query variants, source filtering, one synthesis call, evidence cards, scripts, screenshots, and a visual plan. Do not begin with GPT Researcher's slower multi-step planning flow.
Save the request and response under sources/, and copy the returned report, source list, visual plan, and screenshots into the same topic directory. Use the API defaults in references/local-stack.md.
Apply the returned status strictly:
verified: continue.partial: continue only with the uncertainty stated in the script.not_verified: stop before TTS and explain which core event or premise lacks reliable evidence.
Research current or unstable claims using the package output and primary/major-media sources. Never convert low-quality search snippets into confident narration.
3. Choose and tighten the script
Compare all three scripts against the evidence cards. Select the version with the clearest causal chain and strongest sources, not automatically the most dramatic hook.
Write:
scripts/script-canonical.txt: the exact factual wording used for captions and review.scripts/script-spoken.txt: optional pronunciation-friendly wording for TTS only.scripts/script-review.json: selected version, evidence-card mapping, source IDs, caveats, and estimated duration.
Keep the spoken and canonical versions semantically identical. Do not remove uncertainty language to make a stronger hook. A 60-second script should usually make one central claim, support it with two or three mechanisms, and end with a concise takeaway.
4. Generate the free voice preview
Use the authorized local reference voice with Qwen3-TTS. Generate the narration, MP3, and preview locally, then run Qwen3-ASR verification against the canonical script. Do not upload the voice reference to any external service.
If ASR finds names, numbers, or finance terms wrong, revise script-spoken.txt or regenerate only the affected segment. The canonical script remains the subtitle truth.
5. Pause at the combined human and paid gate
Before any new HeyGen render, present one compact review package:
- chosen script and evidence status;
- 15-second MP3 preview and full narration path;
- presenter portrait path;
- planned digital-human opening and ending durations;
- estimated paid HeyGen seconds and cost;
- screenshots or contact sheet selected for the middle section.
Ask for explicit approval of script, voice, portrait, and the paid render. This is mandatory even when all free stages were automatic. If the user has already approved this exact hash-addressed package in the current production, do not ask again and do not rerender it.
6. Render only the paid presenter footage needed
Use HeyGen for the opening and ending only, normally about 12–15 seconds each. Keep the middle on real source screenshots. Upload only Qwen-generated narration segments, never the local reference recording.
Prefer a single concatenated opening-plus-ending audio render when the existing workflow supports a deterministic split; otherwise render two bookends. Reuse image, audio, and video asset IDs through the SHA-256 registry. Record every paid job under the topic's editable/ directory and the shared HeyGen registry.
Inspect a contact sheet for mouth shape, face, framing, body motion, glitches, and exact segment boundaries before composition.
7. Build the final Remotion timeline
Use the Qwen full narration as the audio master. Mute visual clips where necessary to avoid duplicated audio. Compose:
- Cover visible from frame 0, normally 1–2 seconds.
- Digital-human opening aligned to its matching narration.
- Real news/report screenshots mapped to the claims they support.
- Digital-human conclusion aligned to its matching narration.
- Three-second YouTube-style follow/like end card.
The cover may be designed from an authorized talking-head screenshot and the current topic title. Do not fabricate news imagery. Middle-section evidence visuals must be authentic screenshots captured during research.
Do not fade a cover to or from transparent over a black canvas. Start with the cover fully visible and use a hard cut or a transition with a non-black backing layer.
8. Align captions to speech, not character counts
Use the canonical script as text truth. Determine cue boundaries from the actual Qwen narration using pauses, silence detection, and alignment. Character-proportional timing is preview-only.
Check the first spoken line, every long pause, proper nouns, numbers, the opening/ending joins, and the final caption. Fix any persistent 0.5–1 second lead or lag before final rendering.
9. Validate and deliver
Run code checks, render, and technical validation. Inspect at least the first frame, opening transition, two middle evidence visuals, digital-human ending, final caption, and CTA end card.
The final handoff must include:
final/<topic>-完整口播视频.mp4- final narration MP3 and WAV in
audio/ - canonical and spoken scripts in
scripts/ - research report, sources, screenshots, and visual plan in
sources/ - Remotion job/state and HeyGen state in
editable/ - frame/contact-sheet and validation evidence in
qa/
Report the evidence status, duration, output specs, paid HeyGen seconds, whether assets were reused, and any remaining caveat.
Resume and Revision Rules
Inspect hashes and saved state before running anything. Resume from the earliest invalid stage:
- factual or wording change → research/script, TTS, affected HeyGen segment, captions, render;
- pronunciation-only change → spoken script, TTS, affected HeyGen segment, captions timing, render;
- voice change → TTS onward;
- presenter or lip-sync issue → affected HeyGen segment onward;
- screenshot, cover, subtitle, or CTA change → Remotion only;
- final encoding issue → render/validation only.
Never spend HeyGen credits to fix a local composition or subtitle problem.
Completion Condition
The task is complete only when the final MP4 exists, contains audible narration, passes dimension/duration/audio validation, opens on a non-black cover frame, has captions aligned to speech, uses real evidence visuals, ends with the CTA, and all artifacts live under the single topic directory.
