Back to all companies
Sign in
Back to all companies
Pachyderm logo

Pachyderm

Winter 2015Acquired

Data Versioning, Data Pipelines, and Data Lineage

Save
Pachyderm logo

Pachyderm

Winter 2015Acquired

Data Versioning, Data Pipelines, and Data Lineage

Save
Company details

Pachyderm is a tool for production data pipelines. If you need to chain together data scraping, ingestion, cleaning, munging, wrangling, processing, modeling, and analysis in a sane way, then Pachyderm is for you. If you have an existing set of scripts which do this in an ad-hoc fashion and you're looking for a way to "productionize" them, Pachyderm can make this easy for you.

Location
San Francisco, CA, USA
Founded
2014
Category
Developer Tools
YC Directory Pagepachyderm.com
Founders
  • JD
    Joe Doliner
    Founder/CEO
    LinkedIn
  • JZ
    Joey Zwicker
    Founder
    LinkedIn

Pachyderm is a tool for production data pipelines. If you need to chain together data scraping, ingestion, cleaning, munging, wrangling, processing, modeling, and analysis in a sane way, then Pachyderm is for you. If you have an existing set of scripts which do this in an ad-hoc fashion and you're looking for a way to "productionize" them, Pachyderm can make this easy for you.

Location
San Francisco, CA, USA
Founded
2014
Category
Developer Tools
YC Directory Pagepachyderm.com
Founders
  • JD
    Joe Doliner
    Founder/CEO
    LinkedIn
  • JZ
    Joey Zwicker
    Founder
    LinkedIn

Pressure-test this opportunity

Turn this teardown into a decision-ready prompt for ChatGPT, Claude, or your agent.

On this page
  • Overview
  • Founding Story
  • Timeline
  • What They Built
  • Market Position
  • Target Customers
  • Market Size
  • Competition
  • Business Model
  • Traction
  • Post-Mortem
  • Acquisition as platform consolidation
  • The independent-company countercase
  • What survived
  • Key Lessons
  • Sources

This report was generated by our Deep Research agent and may contain mistakes.

Did we get something wrong? DM @oscrhong and we'll fix it ASAP!

Startups.RIP — Dead startups, alive ideas
PricingContactPrivacyGot feedback? DM @oscrhong
Exec Briefing

Actionable insights

If you only have a few minutes to spare, here’s what investors, operators, and founders should know about Pachyderm (W15).

  1. Reproducibility begins with data state. Pachyderm recognized that code and checkpoints cannot reproduce results without exact inputs and transformations.
  2. Incremental work can be the economic feature. Customer case studies reported large savings by processing changed records rather than complete datasets.
  3. Open source does not solve enterprise distribution. Pachyderm still needed security, management, support, and an expensive route into large accounts.
  4. Strategic value can exceed standalone value. HPE could combine Pachyderm with compute, model development, and existing enterprise relationships.

Overview

Pachyderm built version control and automated pipelines for data. Founded in 2014 by former RethinkDB engineers Joe Doliner and Joey Zwicker, the Winter 2015 YC company paired Git-like data history with containerized transformations, lineage, and incremental recomputation.[1]

The product did not fail. It solved a hard layer of reproducible machine learning, then joined a buyer that could distribute it as part of a larger system. Hewlett Packard Enterprise acquired Pachyderm in January 2023 to connect its data management with supercomputing and machine-learning development products.[2] The structural mechanism was platform consolidation: Pachyderm managed data state and pipelines, while HPE could bundle those capabilities with compute, model development, and enterprise sales.

Founding Story

Doliner and Zwicker met the problem as infrastructure engineers. Both worked at RethinkDB, and Zwicker later worked on Airbnb's data infrastructure before the pair founded Pachyderm.[3]

“Data infrastructure is really hard,” Zwicker wrote in 2015. He captured the second obstacle in another line: “Each company’s needs feel so completely unique.”[3] Teams could reuse databases and containers, yet data-processing systems still became bespoke combinations of storage, schedulers, scripts, and operational knowledge.

Pachyderm's initial answer arrived during the Docker and CoreOS wave. Instead of asking developers to specialize in Hadoop and MapReduce, it let them package ordinary transformation code in containers and run it as a distributed pipeline.[4] The deeper idea was that reproducibility required more than saving source code. A result also depended on the exact input data, every intermediate transformation, and the execution graph.

That insight carried Pachyderm from a developer tool into machine-learning infrastructure. Data versioning, lineage, and incremental processing became the center of the product, while Kubernetes supplied the execution substrate.

Timeline

  • 2014: Doliner and Zwicker founded Pachyderm after working at RethinkDB; Zwicker also brought Airbnb data-infrastructure experience.[1][3]
  • January 2015: Pachyderm launched as an open-source, container-based alternative to Hadoop and MapReduce complexity.[4]
  • 2016: The company released Enterprise 1.0.[5]
  • 2018: Pachyderm reported a $10 million Series A.[5]
  • August 2020: Microsoft's M12 led a $16 million Series B, joined by Decibel, Benchmark, and YC.[6]
  • 2022: A company fact sheet reported another $12 million, bringing total financing above $40 million. HPE Pathfinder also invested.[5][2]
  • January 12, 2023: HPE announced its acquisition of Pachyderm; financial terms were not disclosed.[2]
  • June 2023: HPE marketed the product as HPE Machine Learning Data Management Software.[7]
  • August 2, 2025: EU AI Act obligations for general-purpose model providers began to apply, including technical documentation and public training-content summaries.[13]

What They Built

Pachyderm combined data repositories with pipeline execution. Repositories held versioned data; commits identified exact states. Pipelines watched inputs and automatically ran when data changed. Lineage connected every output to the inputs and transformations that produced it.[8]

Developers packaged arbitrary code in containers. Pachyderm ran those containers on Kubernetes, separating transformation logic from distributed execution and data state.[8] Incremental processing meant a change did not require recomputing an entire dataset. The system could identify affected work and process only the delta.

This model mattered for machine learning because a model checkpoint alone cannot reproduce a result. Teams also need the training data version, preprocessing steps, labels, and dependency history. An integration with Label Studio extended that lineage into versioned data labeling.[9]

The open-core product split adoption from monetization. Community Edition provided core repositories and pipelines. Enterprise added a console, authentication, role-based access, JupyterHub integration, multi-cluster management, support, and unrestricted scale.[10]

Market Position

Target Customers

Pachyderm targeted data and machine-learning teams with large, changing datasets and strict reproducibility needs. Healthcare, autonomous vehicles, and defense were strong fits because lineage and selective recomputation mattered operationally.

Unlock the full Pachyderm teardown

Read the complete post-mortem, the rebuild playbook, and the exact reasons Pachyderm is still worth studying now.

See Pro plans