platoseed
Cloud-native stream processing
Arroyo is a serverless stream processing platform built on top of the open-source Arroyo streaming engine. We allow companies to transform, filter, aggregate, and join their Kafka streams in real-time just by writing SQL. We charge purely based on usage, with no fixed costs or clusters to manage. In 2025, Arroyo was acquired by Cloudflare and now powers Cloudflare Pipelines
Arroyo is a cloud-native stream processing engine that lets users build real-time pipelines using SQL. It emphasizes sub-second results, scalability from micro to mega events, and deployment across modern cloud runtimes, while aiming to be usable without dedicated streaming experts.
Arroyo provides an analytical SQL-based streaming platform that runs as a compact binary and supports deployment on local machines or production environments via Docker or Kubernetes. It processes data in real time with features like sliding/tumbling/session windows, various joins, over 300 SQL functions, exactly-once semantics, and native formats (JSON, Avro, Parquet, text/binary). It includes a web UI for development and monitoring, a REST API for declarative orchestration, and an extensive set of connectors to popular sources and sinks. It is designed to scale from small RAM/CPU footprints to tens of millions of events per second, and to operate across elastic cloud environments with integration options like Fargate and Kubernetes.
Who itβs for: Data teams and engineers needing real-time processing capabilities, especially those seeking to write streaming pipelines in SQL without specialized streaming expertise.
Recent release notes and blog posts indicate active development and feature expansion; acquisition by Cloudflare mentioned in blog posts, suggesting notable traction and corporate interest.
Micah was previously tech lead for streaming compute at Splunk and Lyft, where he built real-time data infra powering Lyft's dynamic pricing, ETA, and safety features. He spends his time rock climbing, playing music, and bringing real-time data to companies that can't hire a streaming infra team.
Build real-time data pipelines in SQL without a real-time platform team
Arroyo builds real-time data pipelines from Kafka by allowing SQL-based streaming queries (filter, aggregate, window, join) with sub-second results. It fully manages infrastructure, autoscale based on data, and charges based on usage with no minimums; targeted at teams needing real-time data processing without managing a real-time platform.
From the original launch (Feb 2023) β may be outdated.

Software that streams data from databases to warehouses in real-time

Simpler and cheaper multi-cloud and cross-platform deployments.