How Kokoro GUI Started
Kokoro GUI is a project that I originally wanted to make in 2022. At the time I didn’t know how to approach the project so I started it as a single file interface for the Kokoro TTS engine, as a starting point to build a more feature rich UI and Backend later on, after a few months I started using the Gemini CLI to make feature updates, but every time I tried to expand the project it couldent handel the task. In August os 2026 I got a Claude subscription, and its way better that Gemini leading to me actually building the text based audio workflow I always wanted.
What I want Kokoro GUI to be
In the future versions of Kokoro GUI I want to expand the workflows to handel
anything from a 24 hour long audiobook, automated audio dubbing, or a tts generator
for another program. I want Kokoro GUI to be the one stop shop, for the most part.
I want to push the frontier of text based editing as an alternative to the traditional
scrub and cut software. I want it to be able to import audio and for that to also act
as an editing surface, that you can edit the transcript and it edits the linked audio.
I would like to even add support for inline Sound FX’s for audio books or podcasts,
like intro-music, or explosion.
The upcomeing features are mostly focusing on the workflow of importing a transcript, audio, or pdf/epub. And quickly have a good enough audio output, without having to manually adjust everything or import the generated audio into another program for editing. One feature is for imported caption files, it will generate the audio and sync it to the caption files timings, allowing for fast dubbing of videos like YouTube’s AI dubbing feature, but with more control. allowing smaller creators to share to a larger audience or educators can upload there lectures in other languages without having to rerecord them, as long as the captions are translated well. I have a lot more ideas in notes scattered around my phone but thats the genral direction I want Kokoro GUI to go, an all in one text based audio editor. With a underlining TTS capability.