“We’re probably uncomfortably honest”: SONALAB on the Benefits and Limitations of Voice AI
AI is reshaping workflows across the media industry—especially wherever content needs to be delivered to different markets, in multiple languages, and across various distribution channels in ever shorter timeframes. Dubbing in particular is widely seen as a bottleneck. This is exactly where Potsdam-based startup SONALAB comes in: the team is developing AI-powered solutions for speech production and dubbing that can be integrated directly into professional production pipelines. With the Fraunhofer-Gesellschaft recently joining as a shareholder, SONALAB has reached an important milestone, with its official market launch set to follow at IBC in Amsterdam this September.
We spoke with the two founders, Alexander Wolf and Sören Hübner, about their roots in Fraunhofer research, the current limits of automation, and the growing acceptance of GenAI in the creative industries.
MTH: Hi Sören, hi Alexander! First things first: if you could choose freely, which series or film would you most like to dub?
Alexander: You’re probably expecting us to say Nolan’s Odyssey or some other current Hollywood blockbuster, but that’s actually not what we’re aiming for right now. To be honest, we still think it’s the wrong use case to have AI dub a feature film or series from start to finish. Technically, because cultural translation, emotional vocal performance, and acting are exactly the areas where models still fall short. And politically within the industry, because the associations reject that approach for understandable reasons. We share that view.
What drives us is the opposite: content that would never get another language version at all. Educational material in a minority language. Accessible versions that were never produced because dubbing them didn’t make economic sense. Documentaries that exist in five languages, but not in fifty others. This is not about replacing existing work; it’s about enabling material that otherwise wouldn’t exist in the first place. Nigeria produces more films every year than the US, yet its cultural reach is still limited in part because of language.
So if we had to choose one example, we’d go with the coolest Nigerian arthouse film of next year.
MTH: Can you briefly walk us through a typical SONALAB dubbing workflow using a real production example?
Sören: We have two kinds of tools for that: first, a web application for automated translation and subtitling, and second, a set of plugins for audio and media editing software. We deliberately address different customer groups: automated “80 percent solutions” when speed matters, and then, for professionals, the option to make precise manual edits directly in familiar audio environments. In the latter case, the audio professional selects a speech clip in Premiere, Pro Tools, Nuendo, or one of many other programs, edits that clip in our plugin, and can then place the result back in exactly the same spot without friction. No exporting, no uploading, no endless format ping-pong. What matters to us in general is that the human stays in the loop. The editor decides, corrects, rejects. We don’t deliver a finished product, we deliver a tool.
For fictional content, for the reasons already mentioned and based on close feedback from the industry, we currently do not see the right use case. And that does not only apply at the level of the voice actors; dialogue scripts are also authorial work. Those are creative processes.
We also believe AI can be especially helpful for ADR in the original language. In that case, the text is already written, the role has been cast, and the voice has been provided with consent. And yet the production still depends on a studio session with someone who may already be tied up on another project. That’s not a creative problem, it’s a scheduling problem, and it arises at the very end of the production chain, where the pressure is highest anyway. That kind of post-work can then be handled directly within the editing software using SONALAB, making the job easier for both sound engineers and voice actors.
MTH: You recently had some good news: the Fraunhofer-Gesellschaft is now officially a shareholder in SONALAB. Many founders would see that not only as funding, but also as a kind of professional seal of approval. How do you see it?
Alexander: Honestly, both — but the second matters more. Fraunhofer does not invest in an idea lightly; it comes after a very thorough technical and commercial review. Going through that process was demanding, and that’s exactly why the outcome means so much to us.
In practical terms, it means three things for us. First: signaling power. When we speak with a broadcaster or a major studio, the first question is always whether the technology really delivers on its promise. Having an institutional shareholder that conducts research in this very field answers that question before the first meeting even begins.
Second: access. We are not just a portfolio company, but also a partner in research and development. That means we continue working with the Spoken Language Processing Group at Fraunhofer IIS and building on research that has been developed there over many years.
Third — and this is easy to underestimate: it creates accountability. A shareholder with that name expects you to work properly, whether it comes to data provenance, compliance, or anything else. That aligns perfectly with what we want to do anyway.
MTH: For many people, Fraunhofer is above all a name associated with a strong reputation in applied research. In fact, you’ve had ties to the organization for quite some time. Where does that connection come from?
Alexander: I spent three years at Fraunhofer IIS working in Next Generation Audio and Voice AI, in product and business development. So I saw every day what kind of research was emerging there in speech synthesis, while also knowing from my time as a sound designer at Bavaria Film and ZDF how studios actually work in practice.
That gap is exactly where SONALAB emerged. On one side, there is cutting-edge technology development. On the other, the media industry is under enormous transformational pressure due to generative AI. The idea was never, “Let’s build a better speech model.” It was, “Let’s bring technology into real-world use.”
MTH: Let’s talk about the core topic: AI-based voice dubbing — and you’re not the only ones working on it. Alongside SONALAB, there are other tools and browser-based applications. What sets you apart in the market?
Alexander: That’s true, there are a lot of them. And most are technically strong. So the difference lies elsewhere.
The first point is integration. Almost all competitors are purely browser-based applications. For a professional post-production operation, that means exporting, uploading, waiting, downloading, importing again, and checking the timing. If you’re dealing with forty takes per episode, any time savings disappear very quickly. We work directly inside the editing software. That may sound trivial, but it is absolutely essential if you want to work fast.
The second point is the trust layer. With us, the origin of every voice is documented, every output carries provenance labeling, and both can be verified retrospectively. Starting August 2 of this year, labeling AI-generated content will become mandatory, and we are not merely meeting the minimum requirement — we built that functionality in from the very beginning and have been expanding it continuously.
And the third point: minority languages. Through our cooperation with Fraunhofer, we are able to train new languages into our model with around 20 hours of high-resolution data. That allows us to support projects in smaller languages, such as Sorbian, which simply do not make economic sense for the major providers. For European minority languages and medium-sized markets, that is extremely exciting and creates additional cultural value. That is something we identify with strongly. The world is not made up only of “major” languages; it is far more culturally diverse than that.
MTH: Dubbing is also about cultural translation — timing, humor, performance. Where do you currently see the limits of automation?
Sören: In more places than most voice AI providers are willing to admit right now. We are probably being uncomfortably honest about that.
What works well today is preserving speaker identity, transferring tone, and achieving clean pronunciation in many languages. It also works for high-volume projects where medium quality is sufficient — documentaries, corporate videos, explainer formats, and the like.
The problems begin with the fine details of emotional vocal performance, which humans train for over many years and which are incredibly difficult to replicate — difficult for another person, and difficult for a machine. What works even less well is cultural translation. A pun that only works in German if you reinvent it completely. The timing of comedy. Knowing when a pause has to last one second longer for the punchline to land. Irony that is not in the text at all, but only in the way something is spoken. And acting in the stricter sense: developing a character over six seasons, with shifts and breaks that are dramaturgically motivated.
That is not simply a data problem that can be solved with more computing power. These are decisions that have to be made by someone who understands the story.
That is why our product is deliberately designed so that those decisions remain with the human. We remove the tedious work — the third timing adjustment, the tenth supporting role with two lines. Not the craft itself.
MTH: In the debate around AI dubbing, promises of efficiency, questions of quality, and concerns about jobs all collide. How do you experience that discussion in conversations with studios, production companies, or creatives?
Alexander: Very differently, depending on who we’re talking to. And the skepticism is justified. With studios and production companies, the tone is pragmatic. They are under pressure that is hardly visible from the outside: simultaneous releases in thirty languages, edits that still change two weeks before release, and post-production is the last step in the chain, so that is where the pressure hits hardest. For them, the question is not whether this is coming, but how.
For voice actors, the concern is real, and we take it seriously. We do not say, “Don’t worry, this won’t affect you,” because that would not be credible. We say: if this technology is coming — and it is — then the crucial question is whether there is a framework around it. Is a voice being used with consent? Is it traceable where it has been used? Is compensation flowing back? Consent is built into the core of our product, including through watermarking and usage reporting.
MTH: You describe your product as “trustworthy, European voice AI” now arriving in professional media production. What exactly does that mean?
Sören: For us, it means three things.
First: data provenance. For every voice we work with, it is documented where it comes from — licensed, consented, and with a clearly defined framework for use. We are currently expanding that inventory systematically. It is the part of our work that takes the longest and looks the least spectacular. But it is the prerequisite for everything else: only if a voice can be assigned to a person can that person also be compensated.
Second: provenance labeling. Every speech output generated by our system carries an inaudible watermark in the signal. That means you can later prove in the final mix whether, and which parts, are AI-generated by us. As already mentioned, Article 50 of the EU AI Act has required machine-readable labeling since August 2. We built this in from the very beginning, so we do not now have to retrofit it.
Third: data processing takes place in Europe. We operate the software on our own servers, which are located near Nuremberg, and no data leaves Germany. On top of that, we are not dependent on American or other online service providers such as AWS. Because we manage the server infrastructure ourselves, we can also offer an on-premise version for customers who need it, where the model runs directly within the customer’s own environment. Their content never leaves their infrastructure.
MTH: Why is a European perspective important?
Alexander: Because it opens up entire market segments to us in the first place: public broadcasters, education, healthcare, and public administration. In those areas, a model trained on undocumented data would never make it through procurement, because no one would sign off on that liability risk.
MTH: A quick look back: you officially founded the company in April. How much time passed between the first idea and the commercial register entry? And what was the biggest challenge along the way?
Alexander: From the first serious thought to the entry in the commercial register in April, it was about two years. From the start of the first funding phase, it was around eighteen months. That may sound long for a startup, but during that time we were not sitting around waiting on an idea — we were building different prototypes with real users. By the time we founded the company, we already had alpha testers, signed letters of intent, and an active pilot project. That was a deliberate decision.
The biggest hurdle was neither product development nor money, but simultaneity. Spinning out of a research organization means you have to resolve equity, licensing, and IP questions properly, while at the same time building customer relationships. Each of those tasks seems critically important and could be a full-time job on its own. Sören once said to me with a smile, “I feel like I have three or four jobs at once.” I can confirm that feeling, and it is probably one of the biggest challenges. Things often simply take longer than you planned. If I had to give one piece of advice, it would be: talk to your customers as early as possible. They are the only real mirror, and ideally they will give you a clear sense of direction. That saves a lot of time.
MTH: You have been part of the MTH Accelerator since 2026. How important is your connection to the regional ecosystem? And at what stage of your company’s development is the program helping you most?
Alexander: The ecosystem is very important for us, above all because of access to the network. Media technology is a field where trust often runs through personal connections. For us, that means potential customers, partners, and of course investors, but also sparring partners and conversations over coffee or dinner. The program comes at exactly the right time for us. We are at the stage where it will be decided how successful our immediate market entry will be. That includes topics such as sales and pricing models, but also contract structures and data processing issues. Those are exactly the questions where an industry network helps.
MTH: Finally — what’s next for SONALAB? And which step is most critical for your further development right now?
Alexander: The next major date is IBC in Amsterdam in September. That is where we will launch publicly, both with the platform and with the trust layer as a standalone product. It is the leading trade show for the broadcast industry, so it is an obvious launch moment for us. We would be delighted if anyone would like to visit us at our stand in the Future Tech Ignite Zone in Hall 14.
At the same time, we will roll out additional platform integrations by the end of the year. As a little sneak peek, I can already tell you that we will also be offering our plugin on the Adobe Exchange marketplace. Compatibility is a crucial sales lever for us, because that is where our users are already looking for tools.
From a technical perspective, we have a huge combined wishlist of possible features and improvements, and of course we will continue working through that steadily over the coming months.
MTH: Thank you very much for the interview, Sören and Alexander!