Complete your first video translation task from scratch, plus detailed guides for all 7 AI tools
YouTube / TikTok / Bilibili link.Each tool is an independent page — submit and download separately, no interference between tasks:
| Tool | Entry | What it does |
|---|---|---|
| Speech to Text | /tools/transcribe | Upload audio/video → transcribe speech → generate timestamped subtitles (SRT/TXT) |
| Subtitle Translation | /tools/subtitle | Upload SRT → one-click translate into 34 languages, timestamps preserved |
| Text to Speech | /tools/tts | Text → AI voiceover with 300+ voices, adjustable speed, MP3 output |
| Video Summarizer | /tools/summarize | Long video/audio → auto-extracted key points and overview |
| Subtitle Remover | /tools/video | Box-select to precisely mask hardcoded subtitles / watermarks |
| Video Clipper | /tools/clip | Cut video segments by start/end timestamps |
| Voice Cloning | /tools/clone | Upload a 10-second voice sample → clone a custom voice → saved permanently |
Need full one-stop processing (transcribe → translate → dub → merge)? Use the "AI Video Processing Workbench" (/workspace/workbench).
The platform uses cloud recognition engines — industry-leading quality, no local models to download:
paraformer-realtime-v2, 50+ languages, auto punctuation and number formatting. Requires a DashScope API Key.Fill in the corresponding key in the "️ Cloud Service Settings" panel of the workbench. Every engine has a "Test Connection" button.
Multiple LLMs supported — fill in the "️ Settings" panel of the workbench:
The default "Auto" mode falls back in the order Kimi → DeepSeek → OpenAI → Baidu — if one fails, the next is used automatically.
speech-02-hd with near-human emotion and quality. Requires MiniMax API Key + GroupId./v1/audio/speech endpoint (302.AI etc.). Requires Base URL + Key + model name.Configured cloud engines are automatically placed at the top of the voice dropdown; if Edge fails, it automatically degrades to a configured cloud engine — tasks never get stuck.
Each voice in the dropdown can be previewed; click Favorite to pin it to the top.
/tools/clone).Perfect for interviews, lectures, movie commentary and other multi-speaker videos.
Pay-as-you-go credits — no prepaid subscription, only pay for what you use:
| Task Type | Billing Basis | Rate |
|---|---|---|
| Speech to Text / Subtitle Translation / Summarize / Subtitle Remover / Clip | Audio/video duration | 1 credit / minute |
| Text to Speech / Subtitle Dubbing | Character count | 1 credit / 200 chars |
| Link download + subtitles | Audio/video duration | 2 credits / minute |
First use of a cloud engine establishes a connection and warms up; afterwards it returns to normal speed. Long videos also naturally take time for recognition and dubbing.
Browser uploads are limited to ≤ 4GB per file (recommended). Link downloads have no such limit.
Yes. Enable " Batch Mode" in the workbench and select multiple videos — they queue and run automatically.
Recognition and translation use the same cloud engines as industry benchmarks (Alibaba Bailian / Volcengine / Tencent + Kimi / DeepSeek LLMs), with subtitle refinement support — output quality is fully comparable.
Recognition and dubbing use official vendor APIs with encrypted transmission. Only the necessary audio/video segments are uploaded for processing; nothing is stored or used for training.
Buy a credit pack in "Account → Recharge". Multiple payment methods supported; credits arrive instantly.