Back to all companies
Sign in
Back to all companies
Speechly logo

Speechly

Winter 2022Acquired

Natural Language Understanding API for Speech that runs on device

Save
Speechly logo

Speechly

Winter 2022Acquired

Natural Language Understanding API for Speech that runs on device

Save
Company details

Speechly makes it cost efficient to understand human conversations, in real-time. We do this by enabling high-accuracy spoken language understanding processed right on the end user’s device.

Hundreds of millions of hours of audio data are generated every single day online. Speechly is the first solution that can affordably turn this audio into actionable insights.

Location
Helsinki, Finland
Founded
2016
Category
Artificial Intelligence
YC profilewww.speechly.com
Founders
  • OS
    Otto Soderlund
    Founder
    X / TwitterLinkedIn
  • HH
    Hannes Heikinheimo
    Founder
    X / TwitterLinkedIn

Speechly makes it cost efficient to understand human conversations, in real-time. We do this by enabling high-accuracy spoken language understanding processed right on the end user’s device.

Hundreds of millions of hours of audio data are generated every single day online. Speechly is the first solution that can affordably turn this audio into actionable insights.

Location
Helsinki, Finland
Founded
2016
Category
Artificial Intelligence
YC profilewww.speechly.com
Founders
  • OS
    Otto Soderlund
    Founder
    X / TwitterLinkedIn
  • HH
    Hannes Heikinheimo
    Founder
    X / TwitterLinkedIn

Pressure-test this opportunity

Explore the risks and possibilities with a prompt for ChatGPT, Claude, or your agent.

On this page
  • Overview
  • Founding Story
  • Timeline
  • What They Built
  • Market Position
  • Target Customers
  • Market Size
  • Competition
  • Business Model
  • Post-Mortem
  • The wedge became more valuable than the platform
  • Moderation rewards owners of the feedback loop
  • Acquisition was an outcome, not proof of a durable standalone business
  • Key Lessons
  • Sources

AI-researched. Check the sources before making a decision.

Found a mistake? Let @oscrhong know.

Startups.RIP — Good ideas. Better timing.
PricingContactPrivacyGot feedback? DM @oscrhong

Speechly (W22) at a glance

  1. The wedge beat the platform. Streaming speech infrastructure became strategically valuable when it addressed live moderation, a buyer risk with an urgent operational owner.
  2. SDK breadth was access, not a moat. Browser, mobile, Unity, and gRPC clients eased adoption, but platform owners controlled the reports, sanctions, appeals, and behavioral context that improved policy decisions.
  3. A good exit can erase the product. Roblox paid $10.118 million for technology, workforce, and synergies; the public API and repositories were later archived. Acquisition value did not prove an enduring standalone franchise.
  4. Known funding is not known economics. Outside capital and purchase consideration are public, while revenue, retention, ownership, and customer concentration are not. Honest analysis stops where those gaps begin.

Overview

Speechly was a Helsinki voice-technology company founded in 2016 by Otto Söderlund and Hannes Heikinheimo. It sold developers a streaming system that combined automatic speech recognition with natural-language understanding, returning tentative and final transcripts, intents, and entities while a person was still speaking. Its SDK footprint stretched from browsers and mobile apps to Unity, but the sharper use case became real-time voice-chat moderation for games and user-generated-content platforms.[1][2]

Speechly did not fail in the conventional sense. Roblox acquired all of its equity on September 18, 2023 for $10.118 million, then the standalone public API and open-source surface disappeared. The deal validates Speechly's technology while exposing the limits of its original position: a general developer API was easier to absorb than defend, while moderation became most valuable inside a platform that owned the conversations, policy labels, and enforcement loop.[3]

Founding Story

Speechly began in Helsinki in 2016. Söderlund had founded and sold a Nordic digital-management consultancy. Heikinheimo brought a machine-learning doctorate and experience spanning Apple Siri, Angry Birds, Nokia Music, and Google Maps. YC lists the company in its Winter 2022 batch and records a team of ten, but the available sources do not explain how the founders met.[1]

The original insight was narrower than the broad promise of a voice assistant. Conventional assistants handled short commands but became frustrating when users issued complex, changing requests. Speechly's response was to combine recognition and language understanding in one streaming interaction, so software could react before the speaker finished. In 2019, Söderlund told TechCrunch, “Voice has shown real promise,” but said current platforms failed on complex requests. He positioned Speechly as a more reactive, multimodal alternative.[4]

That premise led to an API-first company, not a consumer assistant. Speechly supplied the speech layer while customers owned the interface and workflow. Over time, its public materials expanded from cloud clients into mobile, Unity, command-line, and on-device decoder examples. The company then concentrated on live voice moderation, where partial results mattered because intervention after a toxic exchange was already late.

That evolution gave the founders a more concrete mission. At acquisition, Heikinheimo wrote, “Safety and civility are foundational to Roblox,” and said Speechly's AI expertise would address real-time UGC moderation at Roblox's scale.[5] The two remarks bookend the product shift: from making complex commands understandable to making live human conversation governable.

Timeline

  • 2016: Söderlund and Heikinheimo founded Speechly in Helsinki.[1]
  • December 2019: Speechly announced a €2 million seed round led by Cherry Ventures, with Seedcamp, Quantum Angels, Joyance Partners, Social Starts, Tiny.vc, and angels participating.[4]
  • Winter 2022: Speechly joined Y Combinator's W22 batch.[1]
  • September 18, 2023: Roblox acquired all outstanding equity in Speechly, Inc. and Speechly Oy for $10.118 million in total consideration.[3]
  • September 20, 2023: The companies announced the acquisition publicly, with Speechly framed around real-time voice moderation.[5]
  • January 7, 2025: Speechly's GitHub organization was marked archived. Its repositories now state, “We've joined Roblox.”[6]

What They Built

Speechly turned live speech into structured application events. A client opened a bidirectional stream, first sending configuration such as audio encoding, sample rate, channel count, and language. It then sent audio continuously. Speechly returned results tied to an audio-context identifier, with start and stop events dividing a stream into logical segments.[7]

For an application developer, the important output was not merely a transcript. Browser types exposed tentative and final text, entities, intents, segment changes, and connection state. A shopping interface could react to “red shoes under one hundred dollars” as the words arrived. A moderation client could highlight profanity or classify an utterance as offensive before the conversation moved on.[8]

The distribution surface reflected that developer audience. Speechly shipped browser and React packages, Android and iOS clients, a Unity integration, a command-line tool, gRPC definitions, documentation, and demos. Unity mattered because the eventual moderation wedge lived in multiplayer games rather than ordinary web forms. The public monorepo also included Android and iOS decoder examples for on-device transcription. That establishes a hybrid evolution across cloud streaming and local decoding, but it does not prove that production moderation ran entirely on users' devices.[2]

Speechly's differentiation was the coupling of low-latency recognition with application-specific meaning. Generic transcription vendors could return words. General assistants owned an end-user experience. Speechly sat between them, letting a developer build a custom interface around partial results, intents, and entities. The moderation pivot extended the same architecture from understanding what a user wanted to judging whether speech violated policy.

Market Position

Target Customers

The early product targeted application teams building voice-controlled interfaces without staffing an internal speech group. The SDK breadth suggests a developer funnel feeding larger contracts, though historical customer counts and contract sizes are unavailable. By 2023, games and UGC platforms had become the stronger buyer: they carried voice traffic, faced abuse risk, and needed intervention during a live session.

Market Size

No reliable contemporaneous market-size estimate or Speechly revenue figure surfaced. The best demand signal is company-reported research, relayed by TechCrunch, saying roughly 70% of gamers had used voice chat and 72% of those users had encountered a toxic incident.[5] Those figures support prevalence, not willingness to pay. The addressable market was bounded by platforms large enough to need automated moderation but too small, too early, or too constrained to build it internally.

Competition

Speechly competed on two different maps. In general speech infrastructure, cloud platforms and specialist transcription providers had scale, language coverage, and bundled distribution. Open-source and on-device recognition reduced the value of raw transcription. In moderation, named rivals included Modulate's ToxMod and Spectrum Labs.[5]

The decisive axis was not recognition accuracy alone. Moderation quality depended on latency, policy fit, false-positive control, contextual signals, and a stream of labeled outcomes. A vendor could supply inference across platforms, but the platform owner controlled user reports, enforcement actions, appeals, and behavioral context. Roblox later described an in-house voice-safety system processing millions of voice minutes per day, combining audio style, spoken content, machine-labeled training data, human evaluation, and text classification. Roblox reported a 53% reduction in voice-related abuse reports per daily active user after its English rollout.[9]

Roblox does not attribute that system to Speechly, so direct technical lineage is unproven. The strategic pattern is still visible: once moderation became core platform infrastructure, a large owner had both the incentive and the proprietary feedback loop to internalize it.

Business Model

Speechly appears to have sold API access and enterprise deployments, but no reliable historical price card, revenue, gross margin, retention, or customer concentration was found. Its SOC 2 Type II certification before acquisition points toward enterprise procurement rather than a purely hobbyist developer product.[5]

The company raised €2 million in 2019, and TechCrunch later cited PitchBook's estimate of $7.53 million in outside capital.[4][5] Comparing that funding with the $10.118 million acquisition price cannot establish investor returns because ownership, preferences, dilution, debt, and cash at closing are unknown. It does show a modest strategic transaction rather than a blockbuster infrastructure exit.

Post-Mortem

The wedge became more valuable than the platform

Speechly's general API addressed a real interaction problem, but it occupied a thin layer. Customers supplied the interface, workflow, users, and distribution. Cloud and open-source alternatives could attack transcription, while major platforms could own complete assistants. Speechly responded by moving toward moderation, a use case where streaming classification produced an operational outcome rather than another developer primitive.

That move worked well enough to attract Roblox. The acquisition filing allocated $2.8 million to developed technology and $7.536 million to goodwill tied to workforce and synergies. This was not merely an acqui-hire, but neither did Roblox preserve Speechly as an independent platform. Public repositories were archived, and the former API has no verified ongoing standalone presence.[3][6]

Moderation rewards owners of the feedback loop

Speechly could observe audio and return classifications. A game platform could also observe player history, reports, sanctions, appeals, social context, and whether behavior changed after intervention. Those signals compound into better policy models and enforcement. The vendor's cross-platform position offered breadth; the platform owner's closed loop offered depth.

Roblox's later safety publications make that advantage concrete. Its system uses audio style, spoken content, machine-generated labels, human evaluation, and a final text-classification stage. Another Roblox account describes safety and civility as platform-wide systems rather than a detachable feature.[10] Any claim that Speechly directly powered these systems would be inference, not documented fact. The acquisition nevertheless fits the economics of bringing a strategic capability, and the people who built it, inside the owner of the data loop.

Acquisition was an outcome, not proof of a durable standalone business

The strongest counter-narrative is straightforward: Speechly succeeded. It identified a difficult problem, built credible infrastructure, raised outside capital, earned enterprise security credentials, found a sharper market, and sold to the category's dominant customer. That is materially different from shutting down after running out of money. Söderlund's own announcement was explicit: “I look forward to seeing the team continue this important work at Roblox.”[11]

The narrower judgment is that Speechly did not establish a durable independent distribution or data advantage before acquisition. The $10.118 million consideration, including a $5.3 million holdback contingent on post-acquisition conditions, priced technology and integration value without demonstrating a large recurring-revenue franchise. At announcement, terms remained undisclosed; the later SEC filing supplies the authoritative economics.[3]

The public record cannot settle whether selling was the founders' preferred outcome, a response to financing constraints, or the best offer available. It also cannot tie Speechly code to Roblox's later production architecture. What it can establish is the fate of the standalone product: all equity changed hands, public development stopped, and the voice-moderation thesis continued inside a platform with vastly greater scale.

Key Lessons

  • Sell the policy outcome, not the transcript. Speechly's streaming ASR and NLU were technically credible, but moderation connected that machinery to an urgent buyer risk. The valuable unit was safer live conversation, not another text stream.
  • SDK breadth creates access, not defensibility. Browser, mobile, Unity, CLI, and gRPC support reduced adoption friction. They did not give Speechly the proprietary reports, sanctions, appeals, and behavioral context that a platform owner could use to improve policy decisions.
  • An acquisition can validate the wedge and erase the product. Roblox paid for developed technology plus workforce and synergies, while Speechly's standalone public surface was later archived. Builders should distinguish strategic exit value from proof of an enduring independent market.
  • State gaps before converting funding into economics. Speechly's outside capital and purchase price are known; revenue, retention, investor preferences, and customer concentration are not. Any confident return or burn-rate story would exceed the evidence.

Sources

  1. Y Combinator company profile
  2. Speechly archived monorepo
  3. Roblox 2025 Form 10-K
  4. TechCrunch seed announcement
  5. TechCrunch acquisition report
  6. Speechly GitHub organization
  7. Speechly SLU Go API documentation
  8. Speechly browser-client type declarations
  9. Roblox, Deploying ML for Voice Safety
  10. Roblox, Scaling Safety and Civility
  11. Otto Söderlund acquisition announcement