
Humanloop is the LLM evals platform for enterprises.
If you only have a few minutes to spare, here’s what investors, operators, and founders should know about Humanloop (S20).
Humanloop was a 2020 UCL spinout and YC S20 company that helped teams evaluate and improve products built with large language models. It began with tools for expert data annotation, then moved early into prompt management, model evaluation, tracing, and collaboration between engineers and domain experts.[1][2]
Its endpoint was not a conventional startup failure. Humanloop found a real market, reported meaningful production use, and developed expertise that Anthropic wanted. But the horizontal platform became difficult to defend as model vendors and well-funded independent tools expanded into the same evaluation and observability layer. In 2025 Anthropic hired the founders and roughly a dozen colleagues, while the standalone platform was sunset and, according to Anthropic, its assets and intellectual property were not acquired.[3]
Humanloop emerged from University College London's AI Centre. Raza Habib, Peter Hayes, and Jordan Burgess founded the company with UCL professors David Barber and Emine Yilmaz, then joined Y Combinator in 2020.[1] Public sources establish the academic connection but do not clearly document how the three operating founders first met, so that detail remains unresolved.
The first product addressed a constraint in applied machine learning: organizations needed expert-labelled data, but experts were expensive and slow to coordinate. Humanloop's Programmatic product let specialists express rules, apply them across large datasets, and direct attention toward uncertain edge cases. Index Ventures described this as a way to combine expert judgment with active learning rather than force experts to label every example manually.[4]
Foundation models changed that premise. In a 2024 YC interview, Habib described the team's concern plainly: “the biggest risk to us as a business was that these large language models would get really good.”[2] Better models would reduce the need for the annotation workflow Humanloop sold. The founders responded before the threat fully arrived. They tested a product for evaluating model outputs and coordinating technical and non-technical reviewers. Habib said of that validation: “In the end, it took us two days.” The company reported converting ten prospects into paying customers during that test, a founder-reported result rather than an independently audited figure.[2]
That pivot became the defining decision. Humanloop moved from improving training data to helping companies control what happened after they adopted general-purpose models. Early design partners included Duolingo, Gusto, and Vanta. Their feedback pushed the product toward an enterprise workspace where product managers, domain specialists, and engineers could define quality together.[2]
Humanloop became a shared control room for teams shipping LLM features. Engineers could store prompts as versioned files, compare revisions, call multiple model providers through one interface, and trace how a request moved through prompts, retrieval steps, and tools. Product managers and subject-matter experts could review results in a visual editor without working entirely through code.[6]
Evaluation tied those workflows together. Teams assembled datasets of representative inputs, defined evaluators, and ran experiments before release. Evaluators could be deterministic code, another language model acting as a judge, or human review. Quality gates could then enter continuous integration so a prompt or model change that degraded a known scenario would be caught before deployment. Production traces supplied new examples for later tests, connecting incidents back to the evaluation set.[6]
The product evolved far beyond prompt storage. Humanloop added Flows and OpenTelemetry-based tracing, supported eight model providers, and treated retrieval and tool calls as first-class parts of an AI application. In its 2024 retrospective, the company said it shipped weekly, added more than 50 models, handled millions of logs each day across thousands of AI products, and completed four external penetration tests. Those numbers are company-reported, but they show the operational breadth required to sell an evaluation system to enterprises.[7]
Humanloop's difference was organizational as much as technical. It tried to make evaluation a shared discipline: engineers supplied instrumentation and release machinery, while people who understood the task defined what “good” meant. The tension was scope. Once prompt management, traces, experiments, datasets, human annotation, and release gates lived in one product, Humanloop competed across several categories at once.
Humanloop sold to organizations putting generative AI into consequential workflows. Its pricing page addressed engineers, product managers, and domain experts, then paired code-first and visual workflows with enterprise controls including SSO/SAML, role-based access, private-cloud deployment, regional hosting, SOC 2 Type II, HIPAA agreements, and service commitments.[9] This was a sales-led enterprise product, even though a free tier offered two members, 50 evaluation runs, and 10,000 monthly logs.
Named users included Gusto, Filevine, Dixa, FMG, Athena, and Twain. Humanloop's case studies claimed threefold AI deflection and millions of dollars saved at Gusto, six AI products and doubled revenue at Filevine, and threefold product velocity at Dixa. These are vendor marketing claims, not audited results, but the use cases indicate demand from teams where evaluation quality affected support costs, revenue, or release speed.[10]
No reliable public estimate isolates the market Humanloop actually served. Broader generative-AI infrastructure forecasts would overstate the opportunity because spending on models, application development, observability, and evaluation overlaps. Humanloop also did not disclose annual recurring revenue, customer count, retention, churn, or contract values. That absence prevents a defensible bottom-up market calculation.
The stronger demand signal is behavioral. Companies were moving prototypes into production and needed regression testing, human review, and auditability. Humanloop reported hundreds of deployments and millions of daily logs, while customer stories tied evaluation to operational outcomes.[7] The market was real; its boundaries and the share available to an independent horizontal vendor were uncertain.
Humanloop occupied the overlap of prompt management, LLM observability, and evaluation. That exposed it to independent platforms and to model vendors that could move outward from inference into developer tooling. Langfuse now combines traces, prompts, experiments, evaluations, and human annotation in an open-source product that can be self-hosted.[11] Braintrust spans playground experiments, offline evaluation, and continuous production monitoring.[12] LangSmith and Arize occupy adjacent territory, though the bounded research pass did not inspect their current claims deeply enough for detailed comparison.
The decisive axes were workflow coverage, deployment control, and proximity to the model. Humanloop offered broad enterprise collaboration, but an open-source competitor could win buyers that required self-hosting, while a model provider could integrate evaluation with its own APIs, safety research, and enterprise sales. A horizontal vendor had to integrate every provider and application framework while proving that its neutral layer deserved a separate budget.
Humanloop used a freemium-to-enterprise model. The free tier made experimentation cheap; enterprise customers paid for security, governance, deployment options, support, and larger workloads. Customers supplied their own model-provider API keys, so Humanloop sold workflow and control rather than reselling inference.[9]
The company reported $8 million in total funding from YC Continuity, Index Ventures, LocalGlobe, and industry investors. TechCrunch cited PitchBook at $7.91 million across two seed rounds, a difference consistent with rounding or reporting timing.[6][3] Public evidence does not support estimates of burn, margins, or revenue because verified headcount, contracts, and cloud costs are missing. Nor is there a disclosed acquisition price. The available record therefore cannot show whether the product business reached sustainable economics before the team joined Anthropic.
Humanloop reported more than 300 production deployments in 2024, millions of daily logs across thousands of AI products, 50 product releases, and support for more than 50 newly added models.[7] Its case-study index attached concrete outcomes to six customers, including faster product delivery and lower model-training costs.[10]
These signals establish usage, but not commercial scale. Deployment count is not customer count, “AI products” may include multiple products per organization, and case studies select successful examples. No independent source found in the bounded research pass reported ARR, retention, or paid seats. Humanloop had enough credibility and expertise to attract Anthropic, but the public record cannot distinguish a strong product with limited standalone economics from a rapidly growing business whose owners preferred a strategic team transaction.
Humanloop's acquisition announcement said its team was joining Anthropic and customers would transition off the platform.[13] The dated changelog made the consequence explicit: “Following our acquisition, the Humanloop platform will be sunset on September 8th, 2025.”[8] TechCrunch reported that Anthropic took the three founders and around a dozen engineers and researchers, while saying it did not acquire Humanloop's assets or intellectual property.[3]
That structure matters. It suggests the scarce asset was the team's evaluation and enterprise-workflow knowledge, not a software property Anthropic needed to own. Humanloop had spent years learning how companies define model quality, connect domain experts to engineers, and operate evaluation systems under enterprise security requirements. Anthropic could apply that knowledge inside a model and safety organization with far greater distribution.
The pivot away from annotation was correct and impressively early. It also moved Humanloop into a layer whose boundaries were unstable. Prompt storage expanded into experiments; experiments required datasets and judges; production learning required tracing; enterprise deployment required governance and security. The remedy was a broader platform, visible in Humanloop's GA release and weekly shipping cadence.[6][7]
That expansion improved utility but multiplied competitors. Open-source products could bundle similar functions and offer deployment control. Specialized vendors could go deeper on observability or evaluation. Model providers could make their own APIs easier to test and monitor. Humanloop's non-obvious structural problem was that customer value increased when evaluation sat close to every model and every production trace, but capture became harder because no independent vendor controlled either boundary. Supporting more providers made Humanloop useful; it also imposed integration work that the providers themselves did not bear.
There is no sourced founder post-mortem saying Humanloop lost customers, missed revenue targets, or ran out of cash. The company reported substantial usage shortly before the acquisition, and Anthropic publicly valued the team's tooling and evaluation experience.[3] Calling the outcome a failure would outrun the evidence.
The narrower conclusion is more useful. Humanloop successfully anticipated the evaluation category and built credible enterprise capability, yet the eventual transaction did not preserve its product as an independent asset. The team chose, or accepted, an endpoint where customers had to export their data and migrate. Without disclosed terms, revenue, or runway, the motivation cannot be known. The mechanism is still visible: in a consolidating developer-tool layer, specialized talent can command strategic value even when the standalone platform lacks a durable reason to remain separate.