Back to all companies
Sign in
Back to all companies
Pachyderm logo

Pachyderm

Winter 2015Acquired

Data Versioning, Data Pipelines, and Data Lineage

Save
Pachyderm logo

Pachyderm

Winter 2015Acquired

Data Versioning, Data Pipelines, and Data Lineage

Save
Company details

Pachyderm is a tool for production data pipelines. If you need to chain together data scraping, ingestion, cleaning, munging, wrangling, processing, modeling, and analysis in a sane way, then Pachyderm is for you. If you have an existing set of scripts which do this in an ad-hoc fashion and you're looking for a way to "productionize" them, Pachyderm can make this easy for you.

Location
San Francisco, CA, USA
Founded
2014
Category
Developer Tools
YC profilepachyderm.com
Founders
  • JD
    Joe Doliner
    Founder/CEO
    LinkedIn
  • JZ
    Joey Zwicker
    Founder
    LinkedIn

Pachyderm is a tool for production data pipelines. If you need to chain together data scraping, ingestion, cleaning, munging, wrangling, processing, modeling, and analysis in a sane way, then Pachyderm is for you. If you have an existing set of scripts which do this in an ad-hoc fashion and you're looking for a way to "productionize" them, Pachyderm can make this easy for you.

Location
San Francisco, CA, USA
Founded
2014
Category
Developer Tools
YC profilepachyderm.com
Founders
  • JD
    Joe Doliner
    Founder/CEO
    LinkedIn
  • JZ
    Joey Zwicker
    Founder
    LinkedIn

Pressure-test this opportunity

Explore the risks and possibilities with a prompt for ChatGPT, Claude, or your agent.

On this page
  • Overview
  • Founding Story
  • Timeline
  • What They Built
  • Market Position
  • Target Customers
  • Market Size
  • Competition
  • Business Model
  • Traction
  • Post-Mortem
  • Acquisition as platform consolidation
  • The independent-company countercase
  • What survived
  • Key Lessons
  • Sources

AI-researched. Check the sources before making a decision.

Found a mistake? Let @oscrhong know.

Startups.RIP — Good ideas. Better timing.
PricingContactPrivacyGot feedback? DM @oscrhong

Pachyderm (W15) at a glance

  1. Reproducibility begins with data state. Pachyderm recognized that code and checkpoints cannot reproduce results without exact inputs and transformations.
  2. Incremental work can be the economic feature. Customer case studies reported large savings by processing changed records rather than complete datasets.
  3. Open source does not solve enterprise distribution. Pachyderm still needed security, management, support, and an expensive route into large accounts.
  4. Strategic value can exceed standalone value. HPE could combine Pachyderm with compute, model development, and existing enterprise relationships.

Overview

Pachyderm built version control and automated pipelines for data. Founded in 2014 by former RethinkDB engineers Joe Doliner and Joey Zwicker, the Winter 2015 YC company paired Git-like data history with containerized transformations, lineage, and incremental recomputation.[1]

The product did not fail. It solved a hard layer of reproducible machine learning, then joined a buyer that could distribute it as part of a larger system. Hewlett Packard Enterprise acquired Pachyderm in January 2023 to connect its data management with supercomputing and machine-learning development products.[2] The structural mechanism was platform consolidation: Pachyderm managed data state and pipelines, while HPE could bundle those capabilities with compute, model development, and enterprise sales.

Founding Story

Doliner and Zwicker met the problem as infrastructure engineers. Both worked at RethinkDB, and Zwicker later worked on Airbnb's data infrastructure before the pair founded Pachyderm.[3]

“Data infrastructure is really hard,” Zwicker wrote in 2015. He captured the second obstacle in another line: “Each company’s needs feel so completely unique.”[3] Teams could reuse databases and containers, yet data-processing systems still became bespoke combinations of storage, schedulers, scripts, and operational knowledge.

Pachyderm's initial answer arrived during the Docker and CoreOS wave. Instead of asking developers to specialize in Hadoop and MapReduce, it let them package ordinary transformation code in containers and run it as a distributed pipeline.[4] The deeper idea was that reproducibility required more than saving source code. A result also depended on the exact input data, every intermediate transformation, and the execution graph.

That insight carried Pachyderm from a developer tool into machine-learning infrastructure. Data versioning, lineage, and incremental processing became the center of the product, while Kubernetes supplied the execution substrate.

Timeline

  • 2014: Doliner and Zwicker founded Pachyderm after working at RethinkDB; Zwicker also brought Airbnb data-infrastructure experience.[1][3]
  • January 2015: Pachyderm launched as an open-source, container-based alternative to Hadoop and MapReduce complexity.[4]
  • 2016: The company released Enterprise 1.0.[5]
  • 2018: Pachyderm reported a $10 million Series A.[5]
  • August 2020: Microsoft's M12 led a $16 million Series B, joined by Decibel, Benchmark, and YC.[6]
  • 2022: A company fact sheet reported another $12 million, bringing total financing above $40 million. HPE Pathfinder also invested.[5][2]
  • January 12, 2023: HPE announced its acquisition of Pachyderm; financial terms were not disclosed.[2]
  • June 2023: HPE marketed the product as HPE Machine Learning Data Management Software.[7]
  • August 2, 2025: EU AI Act obligations for general-purpose model providers began to apply, including technical documentation and public training-content summaries.[13]

What They Built

Pachyderm combined data repositories with pipeline execution. Repositories held versioned data; commits identified exact states. Pipelines watched inputs and automatically ran when data changed. Lineage connected every output to the inputs and transformations that produced it.[8]

Developers packaged arbitrary code in containers. Pachyderm ran those containers on Kubernetes, separating transformation logic from distributed execution and data state.[8] Incremental processing meant a change did not require recomputing an entire dataset. The system could identify affected work and process only the delta.

This model mattered for machine learning because a model checkpoint alone cannot reproduce a result. Teams also need the training data version, preprocessing steps, labels, and dependency history. An integration with Label Studio extended that lineage into versioned data labeling.[9]

The open-core product split adoption from monetization. Community Edition provided core repositories and pipelines. Enterprise added a console, authentication, role-based access, JupyterHub integration, multi-cluster management, support, and unrestricted scale.[10]

Market Position

Target Customers

Pachyderm targeted data and machine-learning teams with large, changing datasets and strict reproducibility needs. Healthcare, autonomous vehicles, and defense were strong fits because lineage and selective recomputation mattered operationally.

Market Size

No audited ARR, customer count, retention, or conversion data was observed. Company case studies supply workload signals. A confidential healthcare provider reportedly reduced processing and storage needs by 90% and ran ten times faster by processing changed records rather than a full two-terabyte table.[11] Woven Planet reportedly used Pachyderm for petabyte-scale map pipelines and cut processing time by more than 50%.[12] Both results are vendor-published, not independent audits.

Competition

Pachyderm competed inside a crowded infrastructure layer: orchestration tools, data-versioning systems, cloud ML platforms, Kubernetes workflows, and full MLOps suites. Its differentiation was the union of data history, lineage, and automatic incremental execution.

Its constraint was assembly. Customers still needed object storage, Kubernetes, model development, deployment, governance, and enterprise operations. Large platform vendors could sell an integrated stack through existing accounts. HPE already owned enterprise compute and machine-learning development capability, making Pachyderm's data layer more valuable as a portfolio component.

Business Model

Pachyderm used open core. Free software encouraged technical adoption; Enterprise sold security, management, support, multi-cluster controls, and scale.[10]

The company reported more than $40 million raised across a $10 million Series A, $16 million Series B, and $12 million 2022 round.[5] Primary financing records for every round were not observed.

No audited revenue, gross margin, paid-customer count, open-source conversion, retention, or sales-cycle data was found. Enterprise distribution was clearly important: during the Series B, Zwicker said, “M12 and Decibel will prove essential to our expansion.”[6]

Traction

Pachyderm gained named and demanding deployments. Woven Planet used it for autonomous-driving maps, while Lockheed Martin's AI Factory already combined Pachyderm with HPE's machine-learning environment before the acquisition.[12][2]

The confidential healthcare and Woven Planet performance claims indicate technical value at large scale, with the caveat that Pachyderm published both case studies. HPE's 2022 investment followed by acquisition offers stronger evidence that the capability fit enterprise demand. The product continued under an HPE name rather than disappearing.

Post-Mortem

Acquisition as platform consolidation

Pachyderm was not shut down for lack of demand. HPE bought it to add reproducible data pipelines to supercomputing and machine-learning development products.[2] By June 2023, the commercial product had become HPE Machine Learning Data Management Software.[7]

The mechanism was portfolio completion. Pachyderm solved data versioning, lineage, and pipeline execution, but customers still assembled a wider ML system. HPE could connect those capabilities with compute, development tooling, and an enterprise sales channel. Lockheed Martin's existing combined use supplied a concrete preview of that fit.

Open-core infrastructure also requires expensive sales and support. Zwicker's statement that M12 and Decibel would be “essential to our expansion” shows that technical differentiation alone was not enough.[6] HPE could distribute the product into accounts already buying infrastructure.

The independent-company countercase

Open-source infrastructure companies can become durable categories. Pachyderm had a differentiated core, more than $40 million in company-reported financing, and named users handling demanding workloads. Nothing observed proves it could not have remained independent.

The missing evidence matters. No acquisition price, standalone economics, board deliberations, competing offers, or founder rationale for selling was captured. The acquisition may have reflected an attractive strategic offer rather than a ceiling on the business. The defensible conclusion is that integration increased the capability's value, not that independence was impossible.

What survived

HPE retained the product's central ideas: data versioning, lineage, pipelines, and reproducible processing. Rebranding placed them inside a broader machine-learning portfolio. Pachyderm's category thesis survived; the standalone category seller did not.

Key Lessons

  • Reproducibility begins with data state. Pachyderm recognized that source code and model checkpoints are insufficient without exact inputs and transformations.
  • Incremental work can be the economic feature. Processing changed records rather than complete datasets reduced time and infrastructure in company case studies.
  • Open source solves adoption, not enterprise distribution. Pachyderm still needed security, management, support, and a costly route into large accounts.
  • Strategic value can exceed standalone value. HPE could combine Pachyderm with compute, model development, and existing enterprise relationships.

Sources

  1. Y Combinator company profile
  2. HPE acquisition announcement
  3. Zwicker founding essay
  4. TechCrunch launch report
  5. Pachyderm company fact sheet
  6. Series B announcement
  7. HPE ML data management release
  8. Pachyderm architecture paper
  9. Label Studio integration
  10. Pachyderm enterprise fact sheet
  11. Healthcare case study
  12. Woven Planet case study
  13. European Commission GPAI obligations guidance