Can AI turn an off-the-cuff conversation into a finished podcast?
I watched someone record a podcast at their desk. It was casual. A conversation, a camera, and then they were done.
But I found myself wondering what happened after they pressed stop.
Someone still had to find the story, remove the dead parts, clean the audio, structure the episode, package it and publish it.
I'm testing how much of that work frontier AI models can now do.

Recording looked like the easy part.
At a CUB event, I watched the receptionist record a podcast at her desk. It wasn't a carefully scripted studio production. It was off-the-cuff, and that was what made it interesting.
The source material appeared almost effortlessly. The finished piece didn't. Review, editing, structure, audio cleanup, production, packaging and publishing were still ahead.
Much of that work involves judgement. That became the experiment.
The recording can be off-the-cuff.
The finished episode can't feel that way.
The hard part isn't generating audio.
Transcription, summaries, generated voices, music and copy can all be useful ingredients. They don't automatically add up to something worth listening to.
- What's the story?
- What's actually interesting?
- What should disappear?
- What deserves more time?
- Where does the conversation drag?
- What context does the listener need?
- What should they remember afterwards?
Can frontier models move beyond processing the recording and participate meaningfully in editing it?
A polished input would prove very little.
The first source should be ordinary: a rough interview, an unscripted discussion, a meeting or a collection of voice notes.
Repetition, pauses, tangents, unfinished thoughts and weak questions are part of the test. There may be useful ideas buried inside an otherwise unremarkable conversation.
Can the workflow discover the finished piece inside the mess? If the source already sounds like a finished podcast, I've avoided the interesting problem.
What happens after record?
- Record
- Understand
- Edit
- Produce
- Package
- Publish
- Understand
- Transcribe the conversation and identify its themes, arguments, stories and useful moments.
- Edit
- Decide what belongs, what repeats, what drags and what can disappear.
- Structure
- Create a journey for the listener, rather than keeping everything in the order it was said.
- Produce
- Clean the audio, improve pacing and levels, and use music or transitions only where they help.
- Package
- Create the title, description, chapters, show notes and supporting assets.
- Publish
- Bring the episode to a point where its creator can put their name on it.
These are stages I'm testing. They aren't an automated pipeline I've already proved.
Removing silence is automation. Knowing what to keep is editing.
Audio processing is relatively easy to evaluate. Editorial judgement is harder. A model has to recognise the strongest idea, a weak explanation, missing context or a tangent that goes nowhere.
Sometimes the important moment is buried in ordinary conversation. Sometimes one sentence makes several previous minutes unnecessary. Recognising that is the work I want to test.
Clean audio is production.
Knowing what to keep is editing.
Polished shouldn't mean synthetic.
The obvious failure is technically perfect content that no longer sounds like the people who recorded it. Their opinions, humour, hesitation, disagreement and phrasing are part of the source.
I want to reveal the best version of the conversation that actually happened, while preserving what people meant.
Remove the noise.
Keep the people.
Would I publish it?
A completed pipeline isn't the benchmark. Would I willingly put my name on the finished result?

Still from reference.MOV · Supplied input

Still from version1.mp4 · Supplied input
- Story
- Does the episode actually have a point?
- Pacing
- Does it drag?
- Fidelity
- Did the edit preserve what people meant?
- Voice
- Does it still sound like the people involved?
- Production
- Does it feel finished?
- Intervention
- How much manual work was still required?
Awaiting evidence / No scores yet
The episode might only be the beginning.
If the editorial understanding works, the same source could support a shorter cut, chapters, show notes, an article, quotes, social clips, LinkedIn posts or short-form video.
But the first test is simpler: can AI produce one genuinely good episode from a rough conversation? If that doesn't work, the downstream automation doesn't matter.
I haven't proved it yet.
The idea came from watching a simple recording workflow. Now I need to test the other half: take a deliberately rough conversation, let frontier models help understand, edit and produce it, then listen to the result.
If it still needs hours of human editorial work, that's useful evidence. If it produces something I'd genuinely publish, that's much more interesting.
Next test / Off-the-cuff conversation → publishable episode
- Test 01Source → storyStatus / Not run
- Test 02Story → editStatus / Not run
- Test 03Edit → produced episodeStatus / Not run
- Test 04Would I publish it?Status / Not run