S3 Tables adopts Apache Iceberg V3 with native variant, geometry and deletion vectors
On 30 September 2026, Amazon S3 Tables added support for every data type in the Apache Iceberg V3 specification, from deletion vectors to the variant type, geometry and row lineage. V2 tables upgrade atomically, but only Spark 4.0-based engines and the Parquet format can use the new types.
30 September 2026. Amazon S3 Tables announces support for every data type in the Apache Iceberg V3 specification. Deletion vectors, row lineage, the variant type, geometry, geography, nanosecond timestamps: all become native, and V2 tables upgrade to V3 atomically. Apache Spark 4.0. That is the engine requirement to use the new types, which remain limited to the Parquet format. Why it matters: a compliance query that deletes 50,000 rows from a two-billion-row table no longer writes thousands of small delete files, and semi-structured data stops being a JSON string reparsed on every query.
What V2 left on the table
Apache Iceberg has become the open standard for managing large analytics datasets: petabyte-scale tables, schema evolution, hidden partitioning, time travel, all over Parquet files in an object-store data lake. S3 Tables is the storage layer AWS built to keep those tables performant and cost-effective as they grow, with automated compaction, maintenance, replication and tiering.
But teams running Iceberg V2 hit the same walls as data scales. First wall: a compliance request to delete 50,000 records from a two-billion-row table leaves behind positional delete files that slow queries until the next compaction runs. Second wall: semi-structured events land as JSON strings that every query has to parse. Third wall: geospatial coordinates and nanosecond timestamps get encoded as strings or integers. Each workaround adds storage cost, query latency and pipeline code.
Those costs are not theoretical. The positional delete file is a legacy of the V2 model, where every delete had to point out rows one by one; at petabyte scale those files multiply and push compaction further out. Semi-structured JSON, meanwhile, forces every engine to reparse the whole document just to pull out one field — redundant work when that field could have been stored typed at write time.
What V3 delivers
Iceberg V3 attacks all three directly. Deletion vectors replace V2’s positional delete files with a compact binary format: that 50,000-row delete now writes a single vector file instead of thousands of small deletes, dramatically cutting compaction time and delete-file overhead. Row lineage adds the _row_id and _last_updated_sequence_number columns to every record automatically: downstream pipelines can query those fields to find changed rows without scanning the whole table.
The new types let you store semi-structured, geospatial and nanosecond-precision data natively instead of encoding it as strings or integers. The variant type stores semi-structured data in columnar form: on write, the engine shreds variant data into hidden columns and collects statistics; on read, those statistics enable file pruning that cuts I/O sharply versus parsing raw JSON. Add timestamp(tz) for nanosecond precision, geometry and geography for geospatial data, and unknown for columns with no known type.
Under the hood, the variant type earns its performance by shredding semi-structured data into hidden columns at write time and collecting statistics on them. At query time those statistics drive file pruning, so a predicate that touches only one field can skip whole files that do not match — a capability that parsing a raw JSON string simply cannot offer. That is what turns variant from a convenience into a genuine I/O win.
Variant and row lineage in practice
The example AWS gives is a retail analytics team tracking web and mobile events with heterogeneous shapes: a page view has a URL and a duration, a purchase has items and amounts, a search has terms and result counts. With the variant type, all those shapes live in one table with no predefined schema:
CREATE TABLE my_catalog.namespace.clickstream (
event_id bigint,
event_time timestamp,
user_id string,
payload variant
) USING iceberg TBLPROPERTIES ('format-version' = '3'); Reads then skip PARSE_JSON entirely: on Amazon EMR Spark, the variant_get function pulls a typed field straight out of the document.
For deletes, enabling merge-on-read mode writes a small deletion vector instead of rewriting data files, and S3 Tables compaction handles those vectors automatically on the next maintenance cycle. On the incremental side, a query filtering on _last_updated_sequence_number > 42 returns only the rows modified after sequence 42 — your downstream checkpoint becomes a single integer.
This is the part that quietly removes pipeline code. Instead of a change-data-capture layer that diffs snapshots or replays logs, your downstream jobs read _row_id and _last_updated_sequence_number directly and checkpoint the sequence number. The table itself becomes the source of truth for what changed, eliminating a whole class of fragile, hand-rolled incremental logic.
Migrating from V2 to V3
AWS has smoothed the switch. The upgrade is atomic and rewrites no data:
ALTER TABLE my_catalog.namespace.existing_table
SET TBLPROPERTIES ('format-version' = '3'); Existing V2 readers keep working on the upgraded table until you fully adopt V3 features. On the next compaction cycle, S3 Tables removes the old V2 delete files, and new modifications use deletion vectors automatically. Row lineage fields initialize on the first modification after the upgrade. This is a one-way operation — the Iceberg specification does not support downgrading from V3 to V2 — so verify that every engine touching the table supports V3 before you migrate.
Two operational constraints to know. First, the new types require an engine built on Apache Spark 4.0 or later — AWS Glue 6.0 and Amazon EMR 8.1 and above. Second, they are supported only for tables using the Parquet format, not ORC or Avro; and variant, geometry, geography or nanosecond-timestamp columns cannot appear in a table’s compaction sort order.
The migration itself is a checklist, not a project. Confirm the engine version on the Spark 4.0 line, confirm the table format is Parquet, confirm every reader understands V3, then flip the format-version property and let the next compaction cycle absorb the old V2 delete files. Because S3 Tables runs compaction and maintenance for you, there is no window where you have to babysit a rewrite: the service handles the transition on its normal schedule.
What it changes for an operator
First, the win is concrete for anyone managing compliance deletes or semi-structured data: fewer delete files, less JSON to parse, less geospatial encoding code. Second, the V2 → V3 switch is read-compatible during transition — you can migrate the table and let your V2 readers keep running. Third, interoperability stays broad: S3 Tables and AWS Glue Data Catalog both expose the Iceberg REST Catalog (IRC) API, so engines can connect regardless of the catalog endpoint. The feature is available in all regions where S3 Tables is offered, at no charge beyond standard pricing. The net effect is that Iceberg tables on S3 can now follow the V3 spec without forgoing the managed compaction and maintenance that made S3 Tables attractive in the first place.
Verdict
If your pipelines suffer from positional delete files or semi-structured JSON parsing, migrating to V3 is the highest-ROI short-term move: it removes a whole class of latency without re-architecting. If you must preserve non-Spark readers or ORC/Avro formats, stay on V2 until you have verified each engine’s compatibility — the upgrade is one-way. If you are building new analytics lakes, create them directly in V3 on a Spark 4.0 engine: it is now the state of the art for analytics storage on S3, and the cost of entry is zero. AWS’s message fits in one sentence: the V2 workarounds have no reason to exist for anyone who can adopt Spark 4.0.