AWS is expanding Aurora PostgreSQL with DuckDB, allowing users to analyze historical data in Amazon S3 directly from the database. With this, AWS aims to eliminate the need for organizations to first copy data before operational and historical data can be queried together.
According to The Register, the new feature is primarily aimed at environments where current transaction data needs to be combined with large volumes of older data. Until now, this often required data pipelines to transfer information from S3 to a database. Not only does this consume additional storage and computing power, but it also necessitates keeping multiple copies of the same data in sync.
Aurora PostgreSQL can now directly access data stored in Apache Iceberg and Parquet formats. DuckDB is used to perform the analytical processing within Aurora. This allows applications to run queries via existing PostgreSQL interfaces that combine regular Aurora tables with data from a data lake.
DuckDB is quickly gaining a foothold within AWS
The integration follows shortly after AWS’s acquisition of DuckLabs. The company behind DuckDB was acquired last month. DuckDB is an open-source analytical database designed to run directly within other software. This makes the technology suitable for analyzing large datasets without having to set up a separate database system.
AWS now appears to be rapidly assigning DuckDB a broader role within its platform. With Aurora, the technology is primarily intended to eliminate the need for analytical queries to first pass through other services or infrastructure. This reduces the number of network steps and makes it possible to use a single query for both current and historical data.
AWS sees applications in areas such as real-time dashboards and transactions supplemented with historical information. AI agents are also explicitly mentioned as a use case. With such agents, it is not always known in advance what data they will need during a task. Copying all potentially relevant datasets in advance is therefore difficult to scale.
Data remains in S3
During queries, Aurora attempts to read only the necessary data and columns from S3. Frequently used data can also be cached. Developers gain insight into metrics such as the number of rows scanned, the amount of data read from S3, and cache hits.
AWS demonstrates how this works using a financial application. In this example, seven days of recent transactions in Aurora were combined with five years of historical transaction data in a Parquet file on S3. Aurora was able to automatically derive the file’s schema from the metadata.
For applications requiring a response time of a few milliseconds, it remains possible to copy selected data to native Aurora tables. Analytical read queries can also be routed to a read replica, thereby reducing the load on the operational database.
External Iceberg catalogs are also supported. These can be linked to Aurora via AWS Glue, after which PostgreSQL applications can use tables from different catalogs in the same query.
The functionality is available through the aurora_analytics extension and works with Aurora PostgreSQL versions 17.11, 18.6, and later. AWS does not charge a separate fee for this feature. However, using it may result in higher costs for Aurora computing power and for reading data from S3.