Creador Automático de Vídeos

Creador Automático de Vídeos

Proyecto

Creador Automático de Vídeos

Stack técnico

C#, OpenAI/GPT (LLMs), TTS (síntesis de voz) modelos, FFmpeg, Docker, Redis, SQL

Descripción

I built the Automated Video Creator to address a workflow whose complexity is easy to underestimate. A finished video may look like one artifact, but producing it requires a script, scene structure, visual assets, narration, timing, transitions, audio levels, encoding settings, and publication metadata. The project coordinates those elements through a long-running pipeline implemented with C#, language models, text-to-speech services, FFmpeg, Redis, SQL, and Docker. It allowed me to apply my experience in distributed workflows, generative AI, media automation, data modeling, and operational reliability to a process where every stage can be expensive and failures are rarely convenient.

The pipeline begins by turning a topic and creative brief into a structured plan. Rather than asking a model for a complete video in one response, I separate outline creation, section goals, script drafting, scene selection, and narration preparation. Each stage produces a versioned artifact that can be inspected before the next stage proceeds. This design follows the same principles found in agentic AI architecture: decompose the objective, constrain each worker, preserve state, and introduce checkpoints. A structured plan also creates a stable contract between creative reasoning and deterministic media assembly, preventing later rendering code from depending on unstructured model prose.

Script generation is context-aware and iterative. The system tracks the intended audience, tone, length, subject constraints, and continuity between sections. A reviewer can revise a section without regenerating the entire work, and approved text becomes the source for narration and scene timing. My knowledge of LLM workflows and technical writing influenced this separation between drafting and acceptance. Language models are good at generating alternatives, but they can introduce repetition, unsupported statements, or abrupt transitions. The pipeline therefore treats scripts as reviewable content with provenance, not as trusted instructions that automatically proceed to publication.

Narration adds another boundary. Text-to-speech output depends on pronunciation, pacing, voice settings, punctuation, language, and service behavior. I store narration segments and their metadata independently so a problematic passage can be regenerated without losing the surrounding work. Actual audio duration becomes an input to the visual timeline rather than assuming that a word-count estimate is exact. This is an example of feedback-driven architecture: downstream reality corrects upstream planning. The approach draws on my real-time and data-processing experience, where measured state is more reliable than an assumption embedded early in a workflow.

Visual assets are associated with scenes through explicit identifiers and ordering. They may be generated, selected, or transformed, but the renderer receives a deterministic manifest describing what to display and for how long. FFmpeg then performs the repeatable work of combining images, narration, transitions, and encoding parameters. Keeping the final assembly deterministic is important. AI can assist with creative decisions, while FFmpeg executes a precise media contract. This division makes outputs reproducible, simplifies troubleshooting, and avoids allowing a generative component to control low-level commands or file paths without validation.

Redis and SQL support different aspects of workflow state. Queued jobs allow generation and rendering to proceed asynchronously, while persistent records connect the brief, script versions, assets, narration, timeline, attempts, and final output. Work is divided into resumable stages so a failed render or external-service timeout does not require repeating every successful model call. I use bounded retries, idempotent operations where possible, and clear terminal states. These resilience patterns come from cloud architecture and production diagnostics, and they are particularly valuable when tasks consume time, compute resources, or third-party service capacity.

Security controls focus on the authority of each stage. External model and speech credentials remain outside content, file operations are limited to controlled workspaces, generated paths are validated, and FFmpeg receives constructed arguments rather than unchecked text. Review gates prevent the pipeline from publishing generated claims or unsuitable media automatically. My security-engineering and OWASP knowledge informed these trust boundaries, while compliance-oriented architecture encouraged preserving provenance and decisions. Media generation can involve copyrighted, personal, or sensitive material, so the design favors explicit inputs, traceable assets, and human accountability over invisible end-to-end autonomy.

This project brings together C#, LLM orchestration, TTS, FFmpeg, Docker, Redis, SQL, and reliable background processing in a single media system. It also applies ideas from Azure AI, generative AI architecture, autonomous systems, and leadership: define the goal, assign bounded responsibilities, make intermediate work reviewable, and recover intelligently when reality differs from the plan. The result demonstrates that successful AI media automation is primarily an orchestration problem. Quality depends not only on the models, but also on state, timing, contracts, validation, observability, and the ability for a person to guide the work without restarting the entire process.

Pregunta al asistente de Mariojose

Pregúntame sobre Mariojose

Prueba: "¿En qué se especializa Mariojose?" o "Muéstrame sus servicios y trabajos recientes".