Back to all companies
Sign in
Back to all companies
Humanloop logo

Humanloop

Summer 2020Acquired

Humanloop is the LLM evals platform for enterprises.

Save
Humanloop logo

Humanloop

Summer 2020Acquired

Humanloop is the LLM evals platform for enterprises.

Save
Company details

Humanloop is the LLM evals platform for enterprises. Teams at Gusto, Vanta and Duolingo use Humanloop to ship reliable AI products. We enable you to adopt best practices for prompt management, evaluation and observability.

Location
San Francisco, CA, USA
Founded
2020
Category
Generative AI
YC Directory Pagehumanloop.com
Founders
  • RH
    Raza Habib
    Founder
    X / TwitterLinkedIn
  • JB
    Jordan Burgess
    Founder
    X / TwitterLinkedIn
  • PH
    Peter Hayes
    Founder
    LinkedIn

Humanloop is the LLM evals platform for enterprises. Teams at Gusto, Vanta and Duolingo use Humanloop to ship reliable AI products. We enable you to adopt best practices for prompt management, evaluation and observability.

Location
San Francisco, CA, USA
Founded
2020
Category
Generative AI
YC Directory Pagehumanloop.com
Founders
  • RH
    Raza Habib
    Founder
    X / TwitterLinkedIn
  • JB
    Jordan Burgess
    Founder
    X / TwitterLinkedIn
  • PH
    Peter Hayes
    Founder
    LinkedIn

Pressure-test this opportunity

Turn this teardown into a decision-ready prompt for ChatGPT, Claude, or your agent.

On this page
  • Overview
  • Founding Story
  • Timeline
  • What They Built
  • Market Position
  • Target Customers
  • Market Size
  • Competition
  • Business Model
  • Traction
  • Post-Mortem
  • Capability became more valuable than the standalone platform
  • Breadth created strategic exposure
  • The counter-narrative: this was a strategic exit, not evidence of product failure
  • Key Lessons
  • Sources

This report was generated by our Deep Research agent and may contain mistakes.

Did we get something wrong? DM @oscrhong and we'll fix it ASAP!

Startups.RIP — Dead startups, alive ideas
PricingContactPrivacyGot feedback? DM @oscrhong
Exec Briefing

Actionable insights

If you only have a few minutes to spare, here’s what investors, operators, and founders should know about Humanloop (S20).

  1. Kill the premise early. Stronger foundation models threatened the annotation business, so the founders tested evaluation and reportedly found ten paying prospects in two days. Test a technology shift against the business model, not just the feature roadmap.
  2. Design partners defined the product. Duolingo, Gusto, and Vanta helped turn evaluation into a shared workflow for engineers and domain experts. Close customer work sharpened both the product and its enterprise requirements.
  3. Breadth weakened the control point. Prompt management, tracing, datasets, and release gates reinforced one another, but the company controlled neither model distribution nor open-source deployment. More integrations created utility without guaranteeing capture.
  4. Talent value can outlive product value. Anthropic hired the team while the standalone platform closed and customers exported their data. An acquisition may validate hard-won capability without preserving the software business that created it.

Overview

Humanloop was a 2020 UCL spinout and YC S20 company that helped teams evaluate and improve products built with large language models. It began with tools for expert data annotation, then moved early into prompt management, model evaluation, tracing, and collaboration between engineers and domain experts.[1][2]

Its endpoint was not a conventional startup failure. Humanloop found a real market, reported meaningful production use, and developed expertise that Anthropic wanted. But the horizontal platform became difficult to defend as model vendors and well-funded independent tools expanded into the same evaluation and observability layer. In 2025 Anthropic hired the founders and roughly a dozen colleagues, while the standalone platform was sunset and, according to Anthropic, its assets and intellectual property were not acquired.[3]

Founding Story

Humanloop emerged from University College London's AI Centre. Raza Habib, Peter Hayes, and Jordan Burgess founded the company with UCL professors David Barber and Emine Yilmaz, then joined Y Combinator in 2020.[1] Public sources establish the academic connection but do not clearly document how the three operating founders first met, so that detail remains unresolved.

The first product addressed a constraint in applied machine learning: organizations needed expert-labelled data, but experts were expensive and slow to coordinate. Humanloop's Programmatic product let specialists express rules, apply them across large datasets, and direct attention toward uncertain edge cases. Index Ventures described this as a way to combine expert judgment with active learning rather than force experts to label every example manually.[4]

Foundation models changed that premise. In a 2024 YC interview, Habib described the team's concern plainly: “the biggest risk to us as a business was that these large language models would get really good.”[2] Better models would reduce the need for the annotation workflow Humanloop sold. The founders responded before the threat fully arrived. They tested a product for evaluating model outputs and coordinating technical and non-technical reviewers. Habib said of that validation: “In the end, it took us two days.” The company reported converting ten prospects into paying customers during that test, a founder-reported result rather than an independently audited figure.[2]

That pivot became the defining decision. Humanloop moved from improving training data to helping companies control what happened after they adopted general-purpose models. Early design partners included Duolingo, Gusto, and Vanta. Their feedback pushed the product toward an enterprise workspace where product managers, domain specialists, and engineers could define quality together.[2]

Timeline

  • 2020: Humanloop was founded from UCL's AI Centre and joined Y Combinator's Summer 2020 batch.[1]
  • July 2022: Index Ventures announced Humanloop's suite for expert-guided data annotation and active learning.[4]
  • October 2022: Humanloop launched tooling oriented around large language model applications.[5]
  • June 2024: YC published Habib's account of the pivot and Humanloop launched the High Agency podcast as a content channel.[2]
  • November 2024: The company announced general availability of its evaluation and observability platform and said it had raised $8 million in total.[6]
  • December 2024: Humanloop reported 300-plus production deployments, 50 product releases, and millions of daily logs during the year.[7]
  • July 30, 2025: Platform billing stopped ahead of the acquisition transition.[8]
  • August 13, 2025: Humanloop documented the acquisition-related sunset; TechCrunch reported that the founders and roughly a dozen staff were joining Anthropic.[8][3]
  • September 8, 2025: The Humanloop platform was scheduled to become permanently inaccessible.[8]

What They Built

Humanloop became a shared control room for teams shipping LLM features. Engineers could store prompts as versioned files, compare revisions, call multiple model providers through one interface, and trace how a request moved through prompts, retrieval steps, and tools. Product managers and subject-matter experts could review results in a visual editor without working entirely through code.[6]

Evaluation tied those workflows together. Teams assembled datasets of representative inputs, defined evaluators, and ran experiments before release. Evaluators could be deterministic code, another language model acting as a judge, or human review. Quality gates could then enter continuous integration so a prompt or model change that degraded a known scenario would be caught before deployment. Production traces supplied new examples for later tests, connecting incidents back to the evaluation set.[6]

The product evolved far beyond prompt storage. Humanloop added Flows and OpenTelemetry-based tracing, supported eight model providers, and treated retrieval and tool calls as first-class parts of an AI application. In its 2024 retrospective, the company said it shipped weekly, added more than 50 models, handled millions of logs each day across thousands of AI products, and completed four external penetration tests. Those numbers are company-reported, but they show the operational breadth required to sell an evaluation system to enterprises.[7]

Humanloop's difference was organizational as much as technical. It tried to make evaluation a shared discipline: engineers supplied instrumentation and release machinery, while people who understood the task defined what “good” meant. The tension was scope. Once prompt management, traces, experiments, datasets, human annotation, and release gates lived in one product, Humanloop competed across several categories at once.

Market Position

Target Customers

Humanloop sold to organizations putting generative AI into consequential workflows. Its pricing page addressed engineers, product managers, and domain experts, then paired code-first and visual workflows with enterprise controls including SSO/SAML, role-based access, private-cloud deployment, regional hosting, SOC 2 Type II, HIPAA agreements, and service commitments.[9] This was a sales-led enterprise product, even though a free tier offered two members, 50 evaluation runs, and 10,000 monthly logs.

Named users included Gusto, Filevine, Dixa, FMG, Athena, and Twain. Humanloop's case studies claimed threefold AI deflection and millions of dollars saved at Gusto, six AI products and doubled revenue at Filevine, and threefold product velocity at Dixa. These are vendor marketing claims, not audited results, but the use cases indicate demand from teams where evaluation quality affected support costs, revenue, or release speed.[10]

Market Size

No reliable public estimate isolates the market Humanloop actually served. Broader generative-AI infrastructure forecasts would overstate the opportunity because spending on models, application development, observability, and evaluation overlaps. Humanloop also did not disclose annual recurring revenue, customer count, retention, churn, or contract values. That absence prevents a defensible bottom-up market calculation.

The stronger demand signal is behavioral. Companies were moving prototypes into production and needed regression testing, human review, and auditability. Humanloop reported hundreds of deployments and millions of daily logs, while customer stories tied evaluation to operational outcomes.[7] The market was real; its boundaries and the share available to an independent horizontal vendor were uncertain.

Competition

Humanloop occupied the overlap of prompt management, LLM observability, and evaluation. That exposed it to independent platforms and to model vendors that could move outward from inference into developer tooling. Langfuse now combines traces, prompts, experiments, evaluations, and human annotation in an open-source product that can be self-hosted.[11] Braintrust spans playground experiments, offline evaluation, and continuous production monitoring.[12] LangSmith and Arize occupy adjacent territory, though the bounded research pass did not inspect their current claims deeply enough for detailed comparison.

The decisive axes were workflow coverage, deployment control, and proximity to the model. Humanloop offered broad enterprise collaboration, but an open-source competitor could win buyers that required self-hosting, while a model provider could integrate evaluation with its own APIs, safety research, and enterprise sales. A horizontal vendor had to integrate every provider and application framework while proving that its neutral layer deserved a separate budget.

Business Model

Humanloop used a freemium-to-enterprise model. The free tier made experimentation cheap; enterprise customers paid for security, governance, deployment options, support, and larger workloads. Customers supplied their own model-provider API keys, so Humanloop sold workflow and control rather than reselling inference.[9]

The company reported $8 million in total funding from YC Continuity, Index Ventures, LocalGlobe, and industry investors. TechCrunch cited PitchBook at $7.91 million across two seed rounds, a difference consistent with rounding or reporting timing.[6][3] Public evidence does not support estimates of burn, margins, or revenue because verified headcount, contracts, and cloud costs are missing. Nor is there a disclosed acquisition price. The available record therefore cannot show whether the product business reached sustainable economics before the team joined Anthropic.

Traction

Humanloop reported more than 300 production deployments in 2024, millions of daily logs across thousands of AI products, 50 product releases, and support for more than 50 newly added models.[7] Its case-study index attached concrete outcomes to six customers, including faster product delivery and lower model-training costs.[10]

These signals establish usage, but not commercial scale. Deployment count is not customer count, “AI products” may include multiple products per organization, and case studies select successful examples. No independent source found in the bounded research pass reported ARR, retention, or paid seats. Humanloop had enough credibility and expertise to attract Anthropic, but the public record cannot distinguish a strong product with limited standalone economics from a rapidly growing business whose owners preferred a strategic team transaction.

Post-Mortem

Capability became more valuable than the standalone platform

Humanloop's acquisition announcement said its team was joining Anthropic and customers would transition off the platform.[13] The dated changelog made the consequence explicit: “Following our acquisition, the Humanloop platform will be sunset on September 8th, 2025.”[8] TechCrunch reported that Anthropic took the three founders and around a dozen engineers and researchers, while saying it did not acquire Humanloop's assets or intellectual property.[3]

That structure matters. It suggests the scarce asset was the team's evaluation and enterprise-workflow knowledge, not a software property Anthropic needed to own. Humanloop had spent years learning how companies define model quality, connect domain experts to engineers, and operate evaluation systems under enterprise security requirements. Anthropic could apply that knowledge inside a model and safety organization with far greater distribution.

Breadth created strategic exposure

The pivot away from annotation was correct and impressively early. It also moved Humanloop into a layer whose boundaries were unstable. Prompt storage expanded into experiments; experiments required datasets and judges; production learning required tracing; enterprise deployment required governance and security. The remedy was a broader platform, visible in Humanloop's GA release and weekly shipping cadence.[6][7]

That expansion improved utility but multiplied competitors. Open-source products could bundle similar functions and offer deployment control. Specialized vendors could go deeper on observability or evaluation. Model providers could make their own APIs easier to test and monitor. Humanloop's non-obvious structural problem was that customer value increased when evaluation sat close to every model and every production trace, but capture became harder because no independent vendor controlled either boundary. Supporting more providers made Humanloop useful; it also imposed integration work that the providers themselves did not bear.

The counter-narrative: this was a strategic exit, not evidence of product failure

There is no sourced founder post-mortem saying Humanloop lost customers, missed revenue targets, or ran out of cash. The company reported substantial usage shortly before the acquisition, and Anthropic publicly valued the team's tooling and evaluation experience.[3] Calling the outcome a failure would outrun the evidence.

The narrower conclusion is more useful. Humanloop successfully anticipated the evaluation category and built credible enterprise capability, yet the eventual transaction did not preserve its product as an independent asset. The team chose, or accepted, an endpoint where customers had to export their data and migrate. Without disclosed terms, revenue, or runway, the motivation cannot be known. The mechanism is still visible: in a consolidating developer-tool layer, specialized talent can command strategic value even when the standalone platform lacks a durable reason to remain separate.

Key Lessons

  • Humanloop killed its own premise before the market did. The founders saw that stronger foundation models would reduce annotation demand and validated a new evaluation product with ten paying prospects in two days. The transferable lesson is to test the business consequence of a technology shift, not merely add the technology to the existing product.
  • Design partners shaped an enterprise product. Work with Duolingo, Gusto, and Vanta turned evaluation into collaboration across engineering and domain teams. Humanloop's security and deployment features followed the realities of consequential production use, not a generic checklist.
  • Platform breadth can hide a weak control point. Prompt management, traces, datasets, evaluators, and release gates reinforced one another, but Humanloop controlled neither model distribution nor the open-source deployment layer. Each additional integration increased utility without guaranteeing defensibility.
  • An acquisition can validate capability and end the product. Anthropic hired the team, while Humanloop's customers exported data and the platform closed. Operators should separate the value of accumulated expertise from the durability of the software business that produced it.

Sources

  1. UCL Engineering: Humanloop spinout
  2. Y Combinator interview with Raza Habib
  3. TechCrunch: Anthropic nabs Humanloop team
  4. Index Ventures: Humanloop launches expert-guided AI tools
  5. Humanloop: LLM launch
  6. Humanloop: General availability announcement
  7. Humanloop: How Humanloop evolved in 2024
  8. Humanloop changelog: August 2025 sunset
  9. Humanloop pricing
  10. Humanloop customer case studies
  11. Langfuse
  12. Braintrust evaluation documentation
  13. Humanloop acquisition announcement