Build the Studio Conners Audio Pipeline

September 18, 2026

Build the Studio Conners Audio Pipeline

Can we build and validate a reusable ComfyUI workflow for generating original short instrumental audio beds for Studio Conners vertical videos?

What came of the day

We established the working production model for Studio Conners vertical video. We proved a local ComfyUI audio workflow using ACE-Step 1.5, generated and evaluated real soundtrack candidates, and discovered that the finished soundtrack should become the timing anchor for the final edit rather than forcing generated music to match a fully locked cut. The resulting workflow now uses a rough editorial plan first, generates and analyzes the audio, and only then locks the final cut timing and visual-production instructions. The day also changed how video production fits into Studio Conners. Trying to complete a full Obsession and then produce its video at the end of the same day is not a sustainable default. Instead, weekday Obsessions can feed a video queue, with dedicated Saturday production sessions handling audio, ComfyUI visuals, CapCut assembly, masters, and distribution. We defined the beginnings of a private Studio Conners Video Desk to hold that production state and eventually expose it to ChatGPT through MCP. Each finished video will produce three durable masters: the audio master, silent video master, and final distribution-ready video. The actual test video was prepared substantially, but final visual generation, editing, and distribution remain future production work.

The story

We started the day trying to solve what looked like a fairly narrow production problem: how Studio Conners should create original audio for its vertical videos. The immediate goal was to prove a local ComfyUI workflow, understand how the model behaved, and establish a repeatable way to turn a Production Packet into a soundtrack. The first useful breakthrough came from actually generating music instead of continuing to design the system in theory. The first soundtrack was technically successful, but structurally wrong for the video we had planned. It developed too early and faded too soon. That exposed an important flaw in our original assumption: AI-generated music cannot be treated like a deterministic asset that will obediently match a prewritten cut timeline. We generated another candidate with clearer structural guidance and got much closer to what the video needed. Once we analyzed the actual track, the production model changed. The story and rough visual architecture should still come first, but the soundtrack needs to be generated before the final cut timing is locked. The real musical transitions, buildup, payoff, and ending become anchors for the finished edit. Audio therefore became a checkpoint between editorial planning and visual production. We then applied that thinking to a real Studio Conners video based on the first Studio Conners Obsession. We gathered real website screenshots, identified which visuals should remain documentary and which could benefit from image-to-video generation, worked out the production handoff, and clarified that ComfyUI and CapCut should be execution tools rather than places where creative decisions are made. By the end of the day, the bigger operational problem had also become obvious. Producing a polished vertical video after spending an entire day on an Obsession is too much to treat as a quick end-of-day distribution task. A better rhythm is for weekday Obsessions to feed a video queue and for Saturday to become a dedicated Studio Conners production day. That led to the idea of a private Video Desk inside the Studio Conners back office: a durable place for the script, assets, soundtrack candidates, audio analysis, final cut map, ComfyUI instructions, CapCut instructions, masters, provenance, and distribution state. Eventually, MCP could allow ChatGPT to work directly with that same production record. What began as an experiment in AI-generated audio ended as a clearer production system for Studio Conners video as a whole.

Discoveries

  • Audio should anchor the final edit

    A generated soundtrack cannot be assumed to follow the timing imagined in a rough cut plan. The better workflow is to establish the story and approximate beats first, generate the music, analyze the actual musical structure, and then lock the final visual cut points around its real transitions, buildup, payoff, and ending.

  • Audio generation is a planning checkpoint, not just a production task

    The soundtrack sits between editorial planning and final visual production. A video should become Audio Ready first, then the generated track is accepted or regenerated, and only after that should the project advance toward Production Ready.

  • End-of-day video production is not a sustainable default

    A full Obsession already consumes the available creative energy of the day. Trying to close the Obsession, publish the work, and then produce a polished vertical video creates a second substantial creative job. Video production is better treated as its own dedicated production discipline.

  • Saturday can become the Studio Conners video-production day

    Weekday work can feed a video queue while Saturdays handle the production pipeline: choose worthwhile videos, generate and analyze audio, lock cuts, prepare assets, generate visual shots, edit in CapCut, export masters, and distribute. Not every Obsession has to become a video.

  • A Video Desk would solve a real production-state problem

    Too much important production information currently lives temporarily inside chat: scripts, assets, audio candidates, cut timing, prompts, workflow settings, shot status, editing directions, masters, and distribution. A durable Video Desk could become the canonical record for each video and make production much more repeatable.

  • ComfyUI and CapCut should be execution environments, not decision environments

    By the time a project reaches visual production, the creative decisions should already be recorded in the Video Desk. ComfyUI should receive exact generation instructions, and CapCut should receive exact assembly instructions, so production does not require reconstructing the plan from memory.

  • Every finished video should preserve three masters

    The durable output of the production system should be an audio master, a silent video master, and a final distribution-ready video with audio. This keeps the creative components reusable and preserves a clean archival version of the picture.

Results

  • Studio Conners original-audio workflow

    Proved a working local ComfyUI workflow for generating original instrumental soundtracks for Studio Conners vertical videos using ACE-Step 1.5. The test established duration control, prompt structure, export behavior, and a repeatable generation path.

  • Selected soundtrack for “Making Studio Conners real”

    Generated multiple 30-second soundtrack candidates and selected the second candidate after analyzing its actual musical structure. The track demonstrated the new rule that final visual timing should be locked only after the soundtrack has been generated and accepted.

  • Studio Conners Vertical Video Operating Model v0.3

    Defined the revised operating model for Studio Conners vertical video: weekday Obsessions feed a production queue, Saturdays become dedicated video-production days, generated audio acts as the timing checkpoint, and each finished video preserves an audio master, silent video master, and distribution master.

    Open the file
  • Production handoff standard for ComfyUI

    Established the handoff format that should exist before visual production begins: one complete instruction document plus one matching asset package, so ComfyUI can be used as an execution environment rather than a place to reconstruct creative decisions.

    Open the file
  • First real video production plan

    Applied the new workflow to the August 31 Studio Conners work entry, “Making Studio Conners real.” The video now has a selected soundtrack, analyzed musical anchors, real source screenshots, a rough-to-final cut method, and a defined approach for combining documentary assets with limited AI-generated motion.

  • Video Desk direction

    Defined the beginnings of a private Studio Conners Video Desk: a back-office production system that can hold the queue, scripts, assets, soundtrack candidates, audio analysis, cut maps, ComfyUI instructions, CapCut instructions, masters, provenance, and distribution state, with MCP as a future interface for ChatGPT to work directly with the same production record.

Open threads

  • Build the private Video Desk inside the Studio Conners back office and define its production states, records, and asset model.
  • Design the MCP interface so ChatGPT can read and update Video Desk production state directly.
  • Finish the “Making Studio Conners real” vertical video using the new audio-anchored production workflow.
  • Decide the practical rules for the new Saturday video-production cadence, including how many queued videos should be attempted in one session.
  • Define how weekday Obsessions should be queued and prepared for later video production without creating extra end-of-day work.
  • Build the reusable ComfyUI workflow library for audio generation and the different visual shot types Studio Conners actually needs.
  • Define the standard CapCut assembly package so each Production Ready video can be edited from precise instructions without reconstructing decisions.