
Cloud-native stream processing
Explore the risks and possibilities with a prompt for ChatGPT, Claude, or your agent.
Arroyo was a SQL-first stream-processing company founded in 2022 by Micah Wylde and Jackson Newhouse and acquired by Cloudflare on April 10, 2025. Its Rust engine let data teams filter, join, aggregate, and window live events without operating Apache Flink. Arroyo combined an Apache-licensed engine with a managed, usage-based product.[1]
This was a strategic success, not a failure. A two-founder team turned deep streaming experience into an open-source asset that fit Cloudflare's R2, Queues, Workers, and Data Platform. The price, consideration, retention terms, revenue, and acquisition multiple remain undisclosed.
Arroyo names Micah Wylde as cofounder and CEO and Jackson Newhouse as cofounder and CTO. Wylde had led streaming-compute work at Lyft and Splunk; Newhouse spent a decade building high-scale systems at Quantcast.[2] At Lyft, streaming powered pricing, ETAs, safety, and real-time ML but required specialists. Wylde later said, “Arroyo really came out of my frustration trying to use Flink.” He estimated that only two or three of roughly 100 supported engineers became self-sufficient.[3]
Arroyo made SQL the engine's first-class language rather than a thin interface over Flink's Java APIs. The founders paired usability with checkpointing and recovery because they had seen operational failures, including a broadcast-state bug that grew into terabytes of state and required days of recovery.[3]
YC records a 2022 founding date, while Cloudflare says the team started Arroyo in 2023.[1][4] This may distinguish company formation from public product work, but no source resolves it.
Arroyo compiled SQL through Apache DataFusion into distributed, stateful Rust dataflows. It continuously executed registered queries, checkpointed state for recovery, and targeted sub-second results over high-volume streams.[5]
The architecture separated a control plane from workers and used object storage for checkpoints and state.[10] Version 0.5 added exactly-once delivery of Parquet and JSON to object storage, turning Arroyo into a real-time ETL path.[11] The managed service scaled automatically and charged by volume, query complexity, and window size, with no fixed minimum.[1]
Current repository evidence shows roughly 5,000 stars and continued 2026 pull requests.[12][13] Activity does not guarantee indefinite licensing, full Cloudflare parity, or an independent roadmap.
Arroyo targeted teams needing stateful real-time computation without a specialist Flink platform group. Use cases included fraud, real-time ML, trading, IoT, analytics, and ingestion.
No reliable revenue, contract value, retention, or serviceable-market estimate was found. Wylde said dozens of companies ran pipelines, but named customers and volumes were not disclosed.[14]
Arroyo competed with Flink, Spark Streaming, Kafka Streams, Materialize, RisingWave, Tinybird, and ksqlDB.[15] Its advantage was native SQL, recovery, managed scaling, and open-source self-hosting. Its disadvantage was maturity and a small commercial footprint. Cloudflare changed distribution: Arroyo's stateful SQL fit R2 for state, Queues for transit, and Workers for computation.[4]
Arroyo paired open-source self-hosting with a usage-based managed service. Financing evidence is limited: Employbl reports a $500,000 YC pre-seed, but no primary announcement confirms it.[16] No reliable ARR, margins, burn, runway, valuation, or acquisition multiple was found.
Both parties used completed-acquisition language on April 10, 2025, supporting a closed acquisition announcement rather than a pending agreement. Neither disclosed price, cash or stock mix, earn-out, retention package, or a separate signing and legal-close date.[4]
Arroyo shipped eight releases in 2023, attracted a meaningful Show HN discussion, and grew a public repository to roughly 5,000 stars. At acquisition, Wylde claimed dozens of companies used it for ingestion, anti-fraud, IoT, and trading; this was not independently verified.[14]
Post-deal product evidence is stronger. Cloudflare shipped SQL filtering and transformation with exactly-once delivery to R2 as Iceberg tables or Parquet files in September 2025.[8] Current docs and “Validate Arroyo SQL” confirm surviving integration, not full original-code parity.
Arroyo's outcome was a successful acquisition with product follow-through.
The founders attacked a problem they had operated. Native SQL, recovery, serverless scaling, and usage pricing addressed Flink's specialist burden together. That specificity produced both developer interest and an inspectable technical asset.
Public code, releases, architecture, and discussion gave users and Cloudflare evidence beyond a pitch. A small team could demonstrate an engine, ecosystem, and design philosophy before disclosing large commercial metrics.
Cloudflare already owned adjacent primitives. Arroyo supplied the stateful SQL layer. Wylde announced, “I’m incredibly excited to announce that Arroyo has been acquired by Cloudflare,” and said the team would continue broadening access to stream processing.[14]
Integration was staged. Day-one Pipelines handled ingestion; transformations shipped months later. Current product surfaces prove follow-through without proving full code parity or an indefinite roadmap. The counterargument is that the deal may reflect a difficult standalone market because revenue and named customers were not disclosed. Even so, the observable mechanism is successful: an open technical asset fit a scaled buyer's architecture and later shipped inside its platform.