💬 What Captions mode does
Upload your own vertical video and Sleepy adds full motion graphics captions to it, synced to every word that is spoken. Not plain subtitles and not a little bounce on the text: real kinetic typography, animated shapes and transitions, placed so well it looks like an editor spent hours positioning every screen by hand.
Sound effects are added automatically too. A plain talking video with the right sound effects already sounds like a pro edit, and the difference is night and day.
⚖️ Classic or Composite: how to choose
The switch at the top of the Captions tab is the one decision that matters most.
Classic
The motion graphics play on top of your video. Simple, clean and readable, and it still does far more than other caption tools.
The layout is steady: one row size throughout, wrapping or shrinking only when a line is too long. You set the exact box and its height, so you have full control over how it looks.
Best for: screen recordings, busy footage, videos without one clear subject, and when you want the text exactly where you put it.
Composite (behind the subject)
Sleepy masks the subject on every frame and puts the motion graphics behind them, so the person stands in front of the moving text. The look editors rotoscope by hand for hours, from one upload.
The layout is dynamic: every row gets its own size, even within the same screen, with headers and key words flying in bigger. It looks like the work of a high-end editor, the kind realtors pay for their Instagram videos.
Best for: talking heads, creators on camera, products, anything with a clear subject.
Not sure? If there is a person or a clear subject in the shot, start with Composite. For the full story behind it, read the Composite Captions deep dive.
🎯 Mask & Placement (Composite)
In Composite, open Mask & Placement to steer what the engine does. Skip it and Sleepy works it all out from the footage.
- Masked subject: Auto, Human or Not human. Tell it whether the star of the shot is a person or something else, like a product or a pet.
- Text position: Dynamic (Sleepy places every screen) or Fixed (you place it once, and every screen uses that spot). In Fixed, your box sets where the text goes, and Sleepy still lays the text out dynamically inside it, with a different size for every row and flying headers.
- Text layer: with Dynamic, choose Behind if clean (behind the subject whenever it stays readable) or Always behind. With Fixed, choose In front or Behind.
Why Dynamic is the default
Dynamic placement finds the free space in each frame, never covers the face, and follows the subject only when it matters, gently, so it never looks like tracking. Every screen feels composed by an editor. Start with it, and switch to Fixed only when you need a specific spot.
Pro trick: pin it where it already works
Say Dynamic puts almost every screen beautifully beside your subject, but one screen drops low, under the chest. Switch to Fixed, pin the box to that spot beside the subject (a side or a corner), and retry. Sleepy then does its best to keep every screen there, even when that one screen is a tighter fit.
📍 Placing the text yourself
In Classic, and in Composite with Fixed position, a preview of your video appears with the text on it. Drag the text where you want it, and use the size and rotation controls to fine-tune. That position is used for every screen.
- In Classic the box is exact: its size and height are what you get, with one steady row size.
- In Composite Fixed the box is the area: Sleepy lays the text out dynamically inside it.
Keep the face and any important part of the shot clear, and leave a little room at the edges for the platform's own buttons and captions.
🌍 Languages and translation
Captions understand a large set of languages, and the Caption language setting does two jobs:
- Pick the language that is spoken and the transcription gets even more accurate. Auto works too, including videos that mix languages.
- Pick a different language and your captions are translated into it. Spanish video, English captions, in one step.
Supported languages:
📝 Paste the exact script
Already have the words? Paste the exact script that is spoken in the video and the captions sync to it perfectly, with no transcription guesswork. Because Sleepy renders instead of generating, what you provide is exactly what you get.
🎨 The look: presets, animation and borders
- Presets: one-tap looks made for captions: Bold Pop, Clean Line, Neon Glow, Retro Tape, Karaoke (word-by-word highlight) and Liquid Chrome.
- Animation: choose how the text enters and how it keeps moving once it is on screen, plus Text FX.
- Fonts, textures and borders: your own fonts and textures (Starter plan and up), and text borders, including one you design yourself. The brand identity tutorial shows how.
- Blend modes: Normal, Screen, Overlay, Soft Light or Difference, to change how the graphics mix with your footage.
✨ Effects on the subject, the background, or both
Post-process effects are high-quality shaders applied like a pro would, in about a minute. Stack up to 3: deformers like water and liquid glass, glitch, halftone, dither, CMYK print, comic ink, cubes, color grades and film looks like VHS and bloom.
The part nobody else offers: each effect can target the background only, the subject only, or everything, using the same automatic mask. Restyle the world while you stay real, or turn yourself into a comic drawing in a real room. Beat-reactive effects, like Beat Flash, hit on the bass.
📦 Uploads and limits
- Vertical MP4, up to 150MB.
- Up to 30 seconds, or up to 2 minutes on the Agency plan.
- Classic: about a minute once uploaded. Composite up to 30 seconds: 1 to 2 minutes. Long Composite videos can take up to around 6 minutes, still far faster than captioning 30 seconds at a time and stitching the pieces together.
- Retries are faster, because the masking is already done.
✅ Best practices
- Clean, well-lit footage with clear contrast between subject and background gives the sharpest masks.
- Set the spoken language when you know it. One wrong word on screen undermines a perfect result.
- One dominant speaker and little overlapping talk give the cleanest transcription.
- Trim long silent intros and outros before uploading.
- Refine with Retry: render once, then adjust the script, placement or effects and retry. It is fast, and it will not fall apart, because Sleepy renders rather than generates.