Workspace
About ViralReels AI Engine
Discover the underlying models powering our media processing pipeline.
AI Transcription & Transcription fallbacks
We use **OpenAI Whisper** models to perform sub-word level voice transcription on uploads. In case of API rate exhausts, our system fails over to localized processing engines to ensure uninterrupted service.
Vocabulary Analysis & Hook Extraction
The Whisper transcription vocabulary transcript is analyzed by **Llama 3** (Llama-3.3-70b-versatile or Llama-3.1-8b-instant economy model depending on API tokens optimization modes). The analyzer identifies viral hook segments, highlights key quotes, and scores virality index dynamically.
Smart Crop Layout Reframing
Our backend processes video frames utilizing active speaker detection models. It determines the speaker coordinates during dialogues and uses **FFmpeg** filters to stack splits, focus zoom faces, or auto-crop videos into Portrait (9:16), Square (1:1), or Social (4:5) frames dynamically.
High Quality Audio Mixing
We render subtitles directly on top of generated video clips using custom ASS formatting filters. Active words are zoomed and styled dynamically with custom font face families and highlighting layouts matching popular content styles.