Original video in Chinese.
Key Takeaways
- Give Codex the original video and script. It extracts audio, transcribes Chinese speech with Whisper, matches the transcript to the script, removes mistakes, and preserves breathing pauses before exporting FCPXML for a timeline in DaVinci Resolve.
- In my example, more than 50 minutes of footage became a 19-minute episode. The rough cut previously took about two hours by hand; AI completed this stage in minutes, improving the efficiency of the most time-consuming part of my talking-head workflow by more than tenfold in that case.
- I used Perplexity for an initial plan and Codex to build a local project for DaVinci Resolve and Final Cut Pro exchange. Codex with Remotion or HyperFrames can handle animation for B-roll. I have used the project for my courses and shared its archive in the community.
Codex can automate the rough cut of talking-head A-roll: extract the audio, transcribe it with local Whisper, align it with the original script, remove mistakes and repeated takes, preserve pauses of around 0.4 seconds, and export FCPXML. Import it into DaVinci Resolve or Final Cut Pro to obtain a timeline ready for further color work, mixing, and B-roll.
The AI rough-cut workflow
I give Codex the footage and the script. After a short wait, it exports an FCPXML file. Importing it into DaVinci Resolve creates an edited timeline. I then add B-roll and adjust color and sound.
For my talking-head videos, editing A-roll—the portions where I speak to the camera—is the most time-consuming stage. Even with a teleprompter, I make mistakes and repeat takes.
I have to remove large amounts of footage while keeping natural breathing pauses. Without those pauses, transitions between clips sound strange.
For the first episode of my second course, The Big Investment Logic of the AI Industry Chain, I had more than 50 minutes of footage and a final runtime of 19 minutes. Nearly 60% had to go. I published that episode in the newtype community the day before this account.
Previously, I would have spent about two hours on this stage. AI did it in minutes.
Identify mistakes and preserve breathing room
I first described the idea to Perplexity and asked it to research an initial plan. I gave that plan to Codex for feasibility assessment and improvement, then asked Codex to deliver a project that runs locally.
The project extracts audio and uses local Whisper for Chinese transcription. Because I provide a script, it can match the transcript to that reference, remove errors, and leave breathing room at the start and end of each segment. My preset was 0.4 seconds.
It then exports FCPXML. On import, DaVinci Resolve uses the timecodes to locate the corresponding segments in the original video and builds the timeline.
The file worked in the DaVinci Resolve setup I tested and uses Final Cut Pro’s exchange format. I did not test Jianying in person, so I do not include it in the verified scope of this article.
The project can be run manually, but I let Codex operate it. It also built the project and understands how it works, so I see little need to operate the terminal myself.
I shared the project archive in the community. For animated B-roll, Codex can use Remotion or HyperFrames, as I have shown previously.
Test environment, limitations, and acceptance checks
This workflow was first tested in June 2026 and reviewed on August 24, 2026. The test involved Chinese talking-head footage, the matching script, local Whisper, Codex, and DaVinci Resolve, with a 0.4-second pause preset.
Automatic alignment may still fail around proper nouns, repeated sentences, long pauses, or departures from the script. Before delivery, inspect cut points, synchronization, timeline frame rate, and links to the original media. Successful import does not establish that the edit is correct.
For this kind of video, production has become much more convenient. I encourage you to try the workflow.
Primary sources and next steps
- Apple: FCPXML reference
- OpenAI: Codex Skills documentation
- Add animation with Codex and Remotion
- My full division of work when editing video with AI (Chinese)
- Wiki: Agentic Workflow (Chinese)
That is all for this episode. If you want to understand AI, become a more capable independent creator, and meet people with similar interests, join the newtype community. See you next time.
What to check after importing the timeline
Keep references to the original media and preserve an editable timeline. Duplicate the project before importing and inspecting the result; keep the original video intact.
| Item | Passing condition | If it fails |
|---|---|---|
| Media links | Both audio and video resolve to the original files | Relink the media before checking cuts |
| Time base | Frame rate and starting timecode match the footage | Check project settings and export parameters |
| Sentence boundaries | Starts and endings are complete, without clipped words | Widen the cut using the original footage |
| Retakes and repeated sentences | The retained take is semantically complete | Compare it with the script and the full original segment |
| Pauses and synchronization | Transitions sound natural and lip movement matches the audio | Do not treat 0.4 seconds as a fixed rule for every segment |
The original comparison—over 50 minutes of footage, a 19-minute result, and the editing time—belongs to that specific case. It is not a performance promise for every video. The September 27, 2026 addition supplied an acceptance checklist without new editor-version compatibility tests. Successful import still does not establish that every cut is correct.
For preparing materials and scripts, see Wiki: PDF to knowledge notes (Chinese). For the broader creation process, see How to use Codex for content creation (Chinese).