DuckDB 2.0 Alpha Introduces Asynchronous Remote I/O and Redesigned Recursive Query Engine
The upcoming major release of the analytical database brings automated Parquet pre-fetching for cloud storage, column shredding for semi-structured JSON, and index-backed recursive CTEs.

DuckDB is preparing to ship version 2.0 this fall following the release of an initial alpha build, according to a technical evaluation reported by Hacker News. The upcoming release introduces major architectural overhauls to the analytical database engine, led by default asynchronous input/output for remote object storage, an index-backed recursive query engine, and native semi-structured data storage via a new VARIANT data type.
Under DuckDB 1.5.5, remote storage reads executed synchronously across worker threads, causing CPU cores to sit idle while waiting for network downloads from object stores such as Amazon S3. DuckDB 2.0 decouples network transfer from execution by introducing a dedicated background download thread pool that pre-fetches Parquet row groups into an in-memory buffer while worker threads decode data. In benchmarks testing a 2.2-gigabyte Parquet file containing 228 million rows hosted on S3, the decoupled pipeline delivered between two- and three-fold speedups without requiring query modifications.
The pre-fetching system is controlled by the configuration setting read_ahead_depth, which defaults to automatic scaling based on available system threads; setting the parameter to zero restores legacy DuckDB 1.5 behavior. The analysis notes that datasets fragmented into thousands of small files, such as one-megabyte Parquet chunks, do not benefit meaningfully from the change because per-file network round trips remain the primary bottleneck.
The release also rebuilds DuckDB's recursive common table expression (CTE) engine. In previous versions, evaluating recursive queries required scanning the entire underlying table during every iteration step, multiplying execution time by recursion depth. DuckDB 2.0 replaces this loop by reading the source table once and constructing an in-memory lookup index on the parent column, restricting subsequent passes to newly discovered rows. The DuckDB team claims up to a 40-fold speed improvement on graph reachability tasks, with substantial gains on deep hierarchies such as walking thousands of commits in a Git history, though shallow hierarchies like corporate org charts show little difference.
Semi-structured JSON processing receives native support through a new VARIANT type utilizing columnar shredding. When writing data to disk, the engine identifies fields that recur consistently across rows with uniform data types and automatically unpacks them into physical columns within Parquet row groups. Inconsistent or rare fields remain stored in a binary remainder blob. In tests on a five-million-event dataset, storing structured logs as VARIANT reduced disk consumption to approximately one-third of raw JSON strings while providing query performance comparable to static physical columns.
Additional updates in DuckDB 2.0 include table triggers that capture row modifications using transition tables, support for nested schemas, data manipulation language statements within CTEs to enable atomic row migrations, and version 1.0 of the Quack client-server protocol. Managed cloud platform MotherDuck confirmed it plans to support DuckDB 2.0 close to general availability.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.


