
Natural Language Understanding API for Speech that runs on device
Explore the risks and possibilities with a prompt for ChatGPT, Claude, or your agent.
Speechly was a Helsinki voice-technology company founded in 2016 by Otto Söderlund and Hannes Heikinheimo. It sold developers a streaming system that combined automatic speech recognition with natural-language understanding, returning tentative and final transcripts, intents, and entities while a person was still speaking. Its SDK footprint stretched from browsers and mobile apps to Unity, but the sharper use case became real-time voice-chat moderation for games and user-generated-content platforms.[1][2]
Speechly did not fail in the conventional sense. Roblox acquired all of its equity on September 18, 2023 for $10.118 million, then the standalone public API and open-source surface disappeared. The deal validates Speechly's technology while exposing the limits of its original position: a general developer API was easier to absorb than defend, while moderation became most valuable inside a platform that owned the conversations, policy labels, and enforcement loop.[3]
Speechly began in Helsinki in 2016. Söderlund had founded and sold a Nordic digital-management consultancy. Heikinheimo brought a machine-learning doctorate and experience spanning Apple Siri, Angry Birds, Nokia Music, and Google Maps. YC lists the company in its Winter 2022 batch and records a team of ten, but the available sources do not explain how the founders met.[1]
The original insight was narrower than the broad promise of a voice assistant. Conventional assistants handled short commands but became frustrating when users issued complex, changing requests. Speechly's response was to combine recognition and language understanding in one streaming interaction, so software could react before the speaker finished. In 2019, Söderlund told TechCrunch, “Voice has shown real promise,” but said current platforms failed on complex requests. He positioned Speechly as a more reactive, multimodal alternative.[4]
That premise led to an API-first company, not a consumer assistant. Speechly supplied the speech layer while customers owned the interface and workflow. Over time, its public materials expanded from cloud clients into mobile, Unity, command-line, and on-device decoder examples. The company then concentrated on live voice moderation, where partial results mattered because intervention after a toxic exchange was already late.
That evolution gave the founders a more concrete mission. At acquisition, Heikinheimo wrote, “Safety and civility are foundational to Roblox,” and said Speechly's AI expertise would address real-time UGC moderation at Roblox's scale.[5] The two remarks bookend the product shift: from making complex commands understandable to making live human conversation governable.
Speechly turned live speech into structured application events. A client opened a bidirectional stream, first sending configuration such as audio encoding, sample rate, channel count, and language. It then sent audio continuously. Speechly returned results tied to an audio-context identifier, with start and stop events dividing a stream into logical segments.[7]
For an application developer, the important output was not merely a transcript. Browser types exposed tentative and final text, entities, intents, segment changes, and connection state. A shopping interface could react to “red shoes under one hundred dollars” as the words arrived. A moderation client could highlight profanity or classify an utterance as offensive before the conversation moved on.[8]
The distribution surface reflected that developer audience. Speechly shipped browser and React packages, Android and iOS clients, a Unity integration, a command-line tool, gRPC definitions, documentation, and demos. Unity mattered because the eventual moderation wedge lived in multiplayer games rather than ordinary web forms. The public monorepo also included Android and iOS decoder examples for on-device transcription. That establishes a hybrid evolution across cloud streaming and local decoding, but it does not prove that production moderation ran entirely on users' devices.[2]
Speechly's differentiation was the coupling of low-latency recognition with application-specific meaning. Generic transcription vendors could return words. General assistants owned an end-user experience. Speechly sat between them, letting a developer build a custom interface around partial results, intents, and entities. The moderation pivot extended the same architecture from understanding what a user wanted to judging whether speech violated policy.
The early product targeted application teams building voice-controlled interfaces without staffing an internal speech group. The SDK breadth suggests a developer funnel feeding larger contracts, though historical customer counts and contract sizes are unavailable. By 2023, games and UGC platforms had become the stronger buyer: they carried voice traffic, faced abuse risk, and needed intervention during a live session.
No reliable contemporaneous market-size estimate or Speechly revenue figure surfaced. The best demand signal is company-reported research, relayed by TechCrunch, saying roughly 70% of gamers had used voice chat and 72% of those users had encountered a toxic incident.[5] Those figures support prevalence, not willingness to pay. The addressable market was bounded by platforms large enough to need automated moderation but too small, too early, or too constrained to build it internally.
Speechly competed on two different maps. In general speech infrastructure, cloud platforms and specialist transcription providers had scale, language coverage, and bundled distribution. Open-source and on-device recognition reduced the value of raw transcription. In moderation, named rivals included Modulate's ToxMod and Spectrum Labs.[5]
The decisive axis was not recognition accuracy alone. Moderation quality depended on latency, policy fit, false-positive control, contextual signals, and a stream of labeled outcomes. A vendor could supply inference across platforms, but the platform owner controlled user reports, enforcement actions, appeals, and behavioral context. Roblox later described an in-house voice-safety system processing millions of voice minutes per day, combining audio style, spoken content, machine-labeled training data, human evaluation, and text classification. Roblox reported a 53% reduction in voice-related abuse reports per daily active user after its English rollout.[9]
Roblox does not attribute that system to Speechly, so direct technical lineage is unproven. The strategic pattern is still visible: once moderation became core platform infrastructure, a large owner had both the incentive and the proprietary feedback loop to internalize it.
Speechly appears to have sold API access and enterprise deployments, but no reliable historical price card, revenue, gross margin, retention, or customer concentration was found. Its SOC 2 Type II certification before acquisition points toward enterprise procurement rather than a purely hobbyist developer product.[5]
The company raised €2 million in 2019, and TechCrunch later cited PitchBook's estimate of $7.53 million in outside capital.[4][5] Comparing that funding with the $10.118 million acquisition price cannot establish investor returns because ownership, preferences, dilution, debt, and cash at closing are unknown. It does show a modest strategic transaction rather than a blockbuster infrastructure exit.
Speechly's general API addressed a real interaction problem, but it occupied a thin layer. Customers supplied the interface, workflow, users, and distribution. Cloud and open-source alternatives could attack transcription, while major platforms could own complete assistants. Speechly responded by moving toward moderation, a use case where streaming classification produced an operational outcome rather than another developer primitive.
That move worked well enough to attract Roblox. The acquisition filing allocated $2.8 million to developed technology and $7.536 million to goodwill tied to workforce and synergies. This was not merely an acqui-hire, but neither did Roblox preserve Speechly as an independent platform. Public repositories were archived, and the former API has no verified ongoing standalone presence.[3][6]
Speechly could observe audio and return classifications. A game platform could also observe player history, reports, sanctions, appeals, social context, and whether behavior changed after intervention. Those signals compound into better policy models and enforcement. The vendor's cross-platform position offered breadth; the platform owner's closed loop offered depth.
Roblox's later safety publications make that advantage concrete. Its system uses audio style, spoken content, machine-generated labels, human evaluation, and a final text-classification stage. Another Roblox account describes safety and civility as platform-wide systems rather than a detachable feature.[10] Any claim that Speechly directly powered these systems would be inference, not documented fact. The acquisition nevertheless fits the economics of bringing a strategic capability, and the people who built it, inside the owner of the data loop.
The strongest counter-narrative is straightforward: Speechly succeeded. It identified a difficult problem, built credible infrastructure, raised outside capital, earned enterprise security credentials, found a sharper market, and sold to the category's dominant customer. That is materially different from shutting down after running out of money. Söderlund's own announcement was explicit: “I look forward to seeing the team continue this important work at Roblox.”[11]
The narrower judgment is that Speechly did not establish a durable independent distribution or data advantage before acquisition. The $10.118 million consideration, including a $5.3 million holdback contingent on post-acquisition conditions, priced technology and integration value without demonstrating a large recurring-revenue franchise. At announcement, terms remained undisclosed; the later SEC filing supplies the authoritative economics.[3]
The public record cannot settle whether selling was the founders' preferred outcome, a response to financing constraints, or the best offer available. It also cannot tie Speechly code to Roblox's later production architecture. What it can establish is the fate of the standalone product: all equity changed hands, public development stopped, and the voice-moderation thesis continued inside a platform with vastly greater scale.