AppRank icon AppRank
Retour aux classements
Mimika - AI Voice Studio

Mimika - AI Voice Studio

John Saunders

Photo & Video

★ — (0) 4+ 11.9 MB
Voir sur l'App Store

Description

Mimika - AI Voice Studio is a fully native, on-device text-to-speech and voice-changing app for macOS. Speak any text with realistic voices, clone new voices from short samples, swap one speaker's voice in a recording while keeping the music underneath, and export to WAV, AAC, MP3, or back into the original video file. Everything runs locally. No accounts, no subscriptions, no internet required after first launch. WHAT YOU CAN DO • Speak any text with seven built-in voices, or with voices you clone yourself from a 15-second sample • Re-voice any recording - drop an audio or video file, get it back with a different voice while the original timing and pauses stay intact • Speaker isolation - drop a multi-speaker recording, the app identifies who's talking, then you choose per speaker: keep original, silence them, or re-voice in a different voice • Preserve music and ambient sound under re-voiced speech - optional one-time download separates music from voices so background audio survives the swap • Multi-speaker scripts - write dialogue with speaker tags and pause markers, the app renders each speaker in their assigned voice • AI script writer - describe what you want, your local LLM writes the script (works with LM Studio, Ollama, or any OpenAI-compatible local endpoint) • Voice cloning - drop a clean WAV, get a permanent voice you can use anywhere in the app • AI chat with streaming voice and a Metal-rendered audio orb visualizer • Export to WAV, AAC, MP3, or remux back into the original video file with picture preserved exactly ON-DEVICE BY DESIGN Mimika - AI Voice Studio uses Apple's Core ML and MLX frameworks to run two voice engines locally: • Pocket-TTS (Kyutai, 100M params) - sub-second first audio, roughly 3× real-time on Apple Silicon. Great for streaming and long content. • Fish Audio S2 Pro (5B params) - slower but higher quality, excellent for cloning a voice from a 15-second reference clip. Both engines share the same voice library - clone once, use everywhere. GREAT FOR • Content creators who need to swap one speaker's voice in a video • Podcasters cleaning up multi-speaker recordings • Accessibility - reading documents aloud in your preferred voice • Localization - re-voicing existing content while preserving original timing • Anyone who wants high-quality voice tools without a subscription or sending audio to a server RESPONSIBLE USE Voice cloning is intended for your own creative projects, accessibility, and content you have the rights to. Please don't use it to impersonate real people without their consent. REQUIREMENTS • macOS 15 or later • Apple Silicon strongly recommended (M1 or later) • ~500 MB disk for included voice models • Optional 287 MB download for the background-preservation feature • Optional LM Studio or any OpenAI-compatible local endpoint for AI script writing and chat Mimika - AI Voice Studio is free and built entirely in Swift. No telemetry, no analytics, no calls home. CREDITS Mimika - AI Voice Studio builds on excellent open work from several projects: • Pocket-TTS voices and the Mimi audio codec — Kyutai Labs (CC-BY-4.0) • Fish Audio S2 Pro speech model — Fish Audio • Parakeet TDT v3 transcription — NVIDIA NeMo (CC-BY-4.0), via FluidAudio • Hybrid Transformer Demucs background separation — Meta AI (MIT) • Vocos vocoder used for voice enhancement — Charactr Inc. (MIT) • Apple Core ML and MLX frameworks Full attribution and license texts: https://github.com/slaughters85j/pocket-tts-macos/blob/main/ACKNOWLEDGMENTS.md

Nouveautés

Version 1.5.12 · 07/09/2026

Chat threads • Solo and Ensemble conversations are now saved as threads. Pick any one up where you left off from the new sidebar. • Pin, rename, and delete threads. Each thread gets a short auto-written summary. • Start a new chat or a new cast without clearing your history. Reasoning models • Models that think before they answer now work in Solo and Ensemble. Their internal reasoning stays out of the transcript and is never read aloud. • Turn Thinking on or off mid-episode. The change applies on the next turn. Ensemble • Skip your turn when the cast invites you in, instead of waiting out the countdown. • Copy, edit, or delete any line in the transcript. Edits steer where the conversation goes next. • Markdown tables now render properly in both Solo and Ensemble. Settings • Settings and Speaker Isolation now fit on smaller screens. • Load and Eject buttons for your local model, plus an option to load it automatically at launch. • Done closes Settings right away. It no longer waits on a model load. • The app now tells you when a newer version is on the App Store. Fixes • Fixed a number of bugs

Informations

Vendeur
John Saunders
Catégorie
Photo & Video
Version
1.5.12
Nécessite
iOS 15.0+
Taille
11.9 MB
Classification par âge
4+