
Cloud-native stream processing
Turn this teardown into a decision-ready prompt for ChatGPT, Claude, or your agent.
If you only have a few minutes to spare, here’s what investors, operators, and founders should know about Arroyo (W23).
Arroyo was a SQL-first stream-processing company founded in 2022 by Micah Wylde and Jackson Newhouse and acquired by Cloudflare on April 10, 2025. Its Rust engine let data teams filter, join, aggregate, and window live events without operating Apache Flink. Arroyo combined an Apache-licensed engine with a managed, usage-based product.[1]
This was a strategic success, not a failure. A two-founder team turned deep streaming experience into an open-source asset that fit Cloudflare's R2, Queues, Workers, and Data Platform. The price, consideration, retention terms, revenue, and acquisition multiple remain undisclosed.
Arroyo names Micah Wylde as cofounder and CEO and Jackson Newhouse as cofounder and CTO. Wylde had led streaming-compute work at Lyft and Splunk; Newhouse spent a decade building high-scale systems at Quantcast.[2] At Lyft, streaming powered pricing, ETAs, safety, and real-time ML but required specialists. Wylde later said, “Arroyo really came out of my frustration trying to use Flink.” He estimated that only two or three of roughly 100 supported engineers became self-sufficient.[3]
Arroyo made SQL the engine's first-class language rather than a thin interface over Flink's Java APIs. The founders paired usability with checkpointing and recovery because they had seen operational failures, including a broadcast-state bug that grew into terabytes of state and required days of recovery.[3]
YC records a 2022 founding date, while Cloudflare says the team started Arroyo in 2023.[1][4] This may distinguish company formation from public product work, but no source resolves it.
Arroyo compiled SQL through Apache DataFusion into distributed, stateful Rust dataflows. It continuously executed registered queries, checkpointed state for recovery, and targeted sub-second results over high-volume streams.[5]
The architecture separated a control plane from workers and used object storage for checkpoints and state.[10] Version 0.5 added exactly-once delivery of Parquet and JSON to object storage, turning Arroyo into a real-time ETL path.[11] The managed service scaled automatically and charged by volume, query complexity, and window size, with no fixed minimum.[1]
Current repository evidence shows roughly 5,000 stars and continued 2026 pull requests.[12][13] Activity does not guarantee indefinite licensing, full Cloudflare parity, or an independent roadmap.
Arroyo targeted teams needing stateful real-time computation without a specialist Flink platform group. Use cases included fraud, real-time ML, trading, IoT, analytics, and ingestion.
No reliable revenue, contract value, retention, or serviceable-market estimate was found. Wylde said dozens of companies ran pipelines, but named customers and volumes were not disclosed.[14]
Arroyo competed with Flink, Spark Streaming, Kafka Streams, Materialize, RisingWave, Tinybird, and ksqlDB.[15] Its advantage was native SQL, recovery, managed scaling, and open-source self-hosting. Its disadvantage was maturity and a small commercial footprint. Cloudflare changed distribution: Arroyo's stateful SQL fit R2 for state, Queues for transit, and Workers for computation.[4]
Arroyo paired open-source self-hosting with a usage-based managed service. Financing evidence is limited: Employbl reports a $500,000 YC pre-seed, but no primary announcement confirms it.[16] No reliable ARR, margins, burn, runway, valuation, or acquisition multiple was found.
Read the complete post-mortem, the rebuild playbook, and the exact reasons Arroyo is still worth studying now.