Tony Ortiz The Designer logo
Back to Home
Field Notes · Podcast production

I Automated the Tedious Parts of My Podcast. I Still Watch Every Episode.

From ChatGPT descriptions in 2022 to Codex, local scripts, Vector, and Sage: how Tony Ortiz built a podcast process that leaves more room for the conversation.

T
Tony Ortiz
•
I Automated the Tedious Parts of My Podcast. I Still Watch Every Episode.

When a podcast conversation is going right, I can get tunnel vision on my guest. For a moment, I forget about everything outside that conversation. It's magic. It also means I don't retain every word, and I'm not sitting there marking every moment that could become a clip.

I built my production system so I could stay in that moment. Getting there took years of recording, editing, consulting, and figuring out which parts of the job actually needed me.

It started with a YouTube description

I started I Will Not Lose in October 2022. Later that year, when ChatGPT arrived, I immediately saw a use for it: writing YouTube descriptions. I even had a segment called “If You Can't Beat 'Em, Join 'Em” about developments in AI. I thought the technology was going to shake up education and the world around us.

Looking back, those early descriptions were rough. Too many emojis. Too many em dashes. But clean writing helps the person trying to understand what an episode is about, and I never felt a need to hide that I used ChatGPT.

Later, I used examples from years of my own sent emails to help the system follow my voice. I wanted the writing to carry more of the person who was actually speaking. Humanity matters. The tools should help us express it.

The first time I watched my computer edit

At first, I was the only subject, switching manually between two inexpensive cameras in post. When I moved into multicamera conversations, I started using AutoPod. Watching the computer switch to whoever was speaking blew my mind. By my estimate, it saved five or six hours of work per episode.

That subscription still cost money when I wasn't using it. I also took a chance on a lifetime Minvo Pro deal for clipping, and it became a useful part of my toolkit. These products helped me understand what was possible and what I wanted from my own process.

By mid-2026, I was using Codex and local scripts to handle much of that repetitive work myself. I estimate that replacing the recurring multicam and clipping tools saves me about $60 a month. That is a subscription comparison from my own setup, not the total cost of running the system. Models, hardware, storage, maintenance, and my time still count.

A repeatable process came first

Before automating this workflow, I had produced more than 50 full episodes. About 25 of those followed a consistent process I could repeat and keep on schedule. I talked to other podcasters, consulted, and spent years in forums working through the technical decisions.

Every extra microphone, camera, or hour of footage used to create another bottleneck. Even adding timestamps to a description takes time when you are trying to maintain a biweekly release.

My recording choices grew out of that reality. For my remote setup, camcorders let me carry what I need in one bag and record without a computer. I usually choose 1080p because it fits the delivery and keeps transfers manageable. Those are choices for my workflow; another production may have a good reason to capture in 4K.

The audio requirement is firmer: a separate channel for every microphone. Individual tracks give the process a reference for who is speaking and let me level each voice independently. A blended recording removes that control. The system works because the inputs are prepared for it.

From three cameras to a reviewable episode

After recording, I sync the three cameras and the separate audio in Premiere. I treat the audio manually, color the footage, and nest it. Once the timeline is prepared as five aligned tracks with matching start and end points, I run my podedit skill.

The workflow identifies the active speaker and applies camera changes with timing that gives the conversation room to breathe. In my experience, it handles the mechanical switching in minutes. That still feels fun to watch when I remember doing it by hand.

Then I add the intro and outro and watch the episode. If a camera has stopped recording, I prepare a usable fallback shot, such as the wide view or the other speaker, so the cut has something to show instead of black frames. Preparation and recovery are part of production.

The broader system combines OpenAI-assisted work with local scripts and media processing. “Local” describes where those files and processing steps live; it does not mean every model involved runs offline.

The shoe moment is why I still watch

During IWNL 80, Andre Williams stopped to show his shoe. Afterward, there was a pause while he put it back on, and he repeated himself. Watching the episode, I simply cut that section down.

The camera switching could be technically correct through that whole moment. Deciding what belongs in the conversation is a different kind of work. A joke might land awkwardly. A tangent might not fit. Sometimes a small trim makes the whole exchange feel better.

Watching also lets me experience the conversation again. When I'm hosting, my attention is on the guest. During review, I regain the context, notice ideas I want to revisit, and mark a specific moment I might want to shape into a 30-second clip.

“Every episode should be viewed at the very least to see if it's watchable.”

I still aim for conversations around 45 minutes to an hour. Automation hasn't changed that. Audience attention and the conversation set the length. Holding a discussion to that shape is a skill in itself.

One finished episode, more places for it to live

Once the finished episode is exported, I feed it into the content workflow. Vector is where I review the episode package: transcripts, description copy, thumbnail work, short reels, and callbacks to the conversation. Sage is the social review desk where the resulting posts can be considered for approval and the next distribution step.

Vector ingest with all nine outputs selected: clips, transcript, title suggestions, quotes, episode links, show notes, YouTube description, timestamps, and thumbnail
The ingest screen makes the scope clear: all nine export options selected. Clips, copy, chapter markers, and thumbnail work start from the same episode. Select the image to view it full size.

IWNL 79 with Phaze Wun makes this concrete. Its package contains six reels and six text callbacks, covering everything from building a catalog before going viral to the human taste behind creative tools. A conversation can keep working across different formats long after the full episode is finished.

Vector showing the Phaze Wun IWNL 79 package, six artifacts and six reels, with its YouTube description and transcript
Phaze Wun's episode package in Vector: the thumbnail, description, transcript, and six reel exports together in one place. Select the image to view it full size.
Watch the “Perspective AI” review reel from IWNL 80, with its branding and captions in motion. Open the full reel.

In Sage, I can review the copy prepared for each platform and decide what is ready to move forward. The episode themes and media are already there, so I can spend my time on the message and the final creative choices. The cards shown here are drafts waiting for that review.

Sage Review showing an IWNL 80 draft, platform choices, and a human Approve control
Sage Review separates a prepared draft from a human approval. The captured card is still marked ready for review.

What I got back was capacity

I used to spend a minimum of six to eight hours per episode. The show mattered enough that I kept doing it, but I couldn't have managed several shows that way.

Today, on work that fits the process, I can develop an episode in roughly three hours. That is my working estimate, and episode length, preparation, and revisions still affect it. The change makes room for client podcasts while keeping my flagship show alive.

I still procrastinate sometimes. I'm not going to pretend software fixed that. What changed is the amount of work waiting on the other side of a recording. One hour of conversation can become weeks of material instead of one finished video I barely had time to promote.

A system I can offer to other shows

Every client starts with a consultation. Some people need technical advice so they can go build a show themselves. Others want a ten-episode starter with branded motion openers and audio production.

The repeatable process lets me offer production at a rate that makes sense for both of us. Clients can concentrate on creating the show while I handle the tedious production work. They do need to record according to the agreed setup so I can protect the quality that my decade of audio production experience demands.

If that setup doesn't fit, the relationship may stay at consultation, or the project needs a separate manual-production scope. Knowing that before recording saves money and frustration.

Know the work before you automate it

I could spend a lot of time building a live camera-switching agent. For now, I don't need one. A human producer makes intuitive decisions that go beyond an active microphone, and livestreaming is a different production problem. Building an elaborate new capture system around a problem I don't have would be overengineering.

My advice is to learn your process end to end. Repeat it until you know where it breaks. Then choose the work you want the tools to take over.

Costs matter, too. A stack at $150 a month becomes $7,200 over four years, before other expenses. That can be hard to justify for a show without revenue. There are plenty of worthwhile reasons to make a podcast, including simply loving it, but a growing subscription bill doesn't guarantee a good process.

Codex and agents have made building tools around my own needs much more practical. The years spent learning production are what let me tell those tools what good work looks like. I hope taking some of the repetitive work out of the way gives more people room to try.

For me, it means I can keep showing up for the part I loved in the first place: sitting across from someone, getting lost in a conversation, and making a show worth watching.