Install
openclaw skills install @social-media-skills/talking-head-and-piece-to-cameraThe on-camera delivery craft — helping a real human film themselves talking to a lens and look like themselves doing it. Use when someone wants a "talking head video" or "piece to camera," says "film myself" or "I look stiff on camera," asks about a teleprompter, framing, lighting, audio, or retakes, or wants to batch-film videos. Uses the TAKES framework. Phone-first: gear is almost never the bottleneck. Reads brand-profile + voice-builder first; takes its script from short-form-video-script (that writes it, this delivers it). The agent coaches setup + delivery, formats prompter/beat-map scripts, and plans batch days; the HUMAN films and picks the take (the agent cannot see footage); WoopSocial publishes the finished file. Camera-shy? Route honestly to heygen/synthesia or faceless formats. Never fabricates "that take looks great." Distinct from scripting-and-storyboarding (the shoot plan), heygen/synthesia (avatars), and captions-and-clipping/capcut/descript (the edit).
openclaw skills install @social-media-skills/talking-head-and-piece-to-cameraThe on-camera delivery craft — tape the setup, anchor the map (not the lines), kick the first 3 seconds,
embrace the retake rules, stack the batch. The script comes from short-form-video-script; the human
films and picks the take; WoopSocial publishes the finished file.
A talking head works because a real face builds parasocial trust an avatar can't (that's exactly why synthesia
routes trust-led founder content here). Three truths most first-timers get backwards. First, gear is not the
bottleneck — a phone at eye level, facing a window, with a cheap lav mic outperforms an expensive camera set up
wrong; viewers forgive soft video and never forgive bad audio. Second, reading kills it — memorize the map
(the beats), not the lines; a word-for-word read shows in the eyes, and a slightly imperfect riff reads as human.
Third, the good-enough take ships — take 4 is usually worse than take 2 because energy decays faster than
delivery improves; perfectionism is a retention strategy for exactly nobody. Deliver 20% more energy than feels
natural, talk to one person, and publish the take where you sound like yourself.
(Depth: references/the-takes-framework.md.)
descript Eye Contact patches a read, not a performance.short-form-video-script; ship the good-enough take.Any recent phone shoots 4K that out-resolves every social feed; audio drives perceived quality more than image
(creator consensus — attribute); a below-eye lens reads as looming, backlit windows silhouette you; on-camera
energy reads ~20% flatter than it feels (broadcast coaching convention); take quality typically peaks by take
2–3 then decays with energy; batch sessions fade after ~60–90 minutes — directional, attribute,
verify-quarterly. Full figures + phone-first setup specifics: references/talking-head-2026-reality.md.
Batch-day recipe, setup recipes (desk / walking / car), and camera-shy on-ramps:
references/batch-filming-and-recipes.md.
references/scope-and-connections.md.)heygen (creator/social lane) or synthesia (enterprise/
L&D lane) for a disclosed avatar, or to faceless formats (screen-record / B-roll + ai-voiceover). Offer the
gentle on-ramp — voice-only first, then hands/desk shots, then face — but never pressure.talking-head-and-piece-to-camera (this) = the human filming/delivery craft · short-form-video-script = the script this delivers (pair) · scripting-and-storyboarding = the multi-scene shoot plan (this is the shoot-day performance) · heygen / synthesia = synthetic presenters when the human can't/won't film · captions-and-clipping / capcut / descript = the edit after the shoot (descript's Eye Contact patches a read; it doesn't replace delivery) · livestream-and-realtime = live to-camera (no retakes) · ai-voiceover = voice without a face.
Reads first: brand-profile + voice-builder. Takes the script from: short-form-video-script (or youtube-long-form), the plan from scripting-and-storyboarding, batch slots from batch-content-plan + content-calendar. Feeds: captions-and-clipping / capcut / descript (the edit), opus-clip (clipping long pieces), cross-platform-repurposing. Routes away: avatars → heygen / synthesia. Publishes via: edited file → scheduling-and-queue → WoopSocial. Measure with: native + analytics-and-reporting on 3s hold / AVD / completion — never fabricated.
A filmed piece to camera delivered from a beat map (first + last lines verbatim, middle riffed), shot phone-first at eye level facing the light with clean close audio and a caption-safe 9:16 frame, opening mid-energy on the hook with no wind-up, retaken per beat under the three-strike rule and shipped at good-enough rather than sanded lifeless, batched (4–8 scripts, top swaps, ≤90 min) when volume is the goal; camera-shy humans routed honestly to heygen/synthesia or faceless formats; the human filmed and picked the take (the agent never judged footage it can't see, never fabricated praise, never prescribed gear as the fix); consent handled for anyone else in frame; the file edited via captions-and-clipping/capcut/descript and published via scheduling-and-queue → WoopSocial; measured on 3s hold / AVD / completion; and correctly distinguished from short-form-video-script, scripting-and-storyboarding, heygen/synthesia, and the editing skills.