Apache Iceberg Go 0.7.0 Release
The Apache Iceberg community is pleased to announce version 0.7.0 of iceberg-go.
This release covers approximately four months of development since the 0.6.0 release in May 2026 and is the result of merging nearly 750 PRs from 53 contributors, including 33 first-time contributors. See the full changelog for the complete list of changes.
iceberg-go is a native Go implementation of the Apache Iceberg table format, providing libraries for reading, writing, and managing Iceberg tables in Go applications.
Release Highlights🔗
Iceberg V3 Table Spec Support🔗
This release continues to close the gap with the Iceberg V3 table specification:
- Geometry and geography types:
geometryandgeographytypes with schema plumbing, WKB encoding and GeoArrow conversion, PROJJSON CRS read/write support, geo bounds emitted in manifest data files and aggregated at the manifest level, and bounding-box predicate evaluation for geo pruning - Variant shredding: A shredded variant reader and an opt-in shredded variant writer with type-uniformity inference, bounds for shredded fields, and predicate pushdown over shredded fields using a new
extractexpression - Deletion vectors: The scanner deletion-vector read path, DV writes wired into the position-delete producer, DVs for partitioned V3 tables, merging DVs on repeated merge-on-read deletes, DV planning validations matching Java's
DeleteFileIndex, and cross-client DV format compliance tests - Row lineage: Row lineage preserved through copy-on-write rewrites,
first_row_idinherited for V1/V2-era manifests in V3 tables, and row-group pruning re-enabled with correct lineage tracking - Unknown transforms: Unknown partition and sort transforms can now be read for forward compatibility
- Nanosecond timestamps: Nanosecond support across predicates, year/month and bucket transforms, Substrait literals, and partition values
Metadata Tables and Incremental Scans🔗
Table.Inspect gained a full set of metadata tables, and two new incremental scan types were added:
- Metadata tables:
history,snapshots,metadata_log_entries,refs,manifests,data_files,delete_files,entries,partitions,position_deletes,files,all_files,all_data_files, andall_delete_files,all_manifests, andall_entries - Incremental append scan:
Table.NewIncrementalAppendScanplans data files added between two snapshots - Incremental changelog scan:
Table.NewIncrementalChangelogScanproduces file-level insert and delete tasks, along with delete-file-backed changelog task types as groundwork for row-level changelogs
REST Server-Side Scan Planning🔗
iceberg-go can now delegate scan planning to a REST catalog that supports the scan-planning endpoints. Local planning remains the default; remote planning is opt-in:
- Foundations: Iceberg expression JSON parsing and the public scan-planning API surface
- Client: Scan-planning client methods, a
WaitForPlanpoller, and decoding of REST scan tasks intoFileScanTasks - Delegation: Complete server-side scan delegation, concurrent fetching of remote plan tasks, and retries, telemetry, and documentation
- Correctness: A Java scan-planning integration suite and verification that REST planning produces tasks equivalent to local planning
Metrics Reporting🔗
A new metrics package adds opt-in reporting for scans and commits:
- Reporter framework: A
Reportercontract with built-in reporters, report and wire types, and reporter selection through catalog, table, and scan - Instrumentation:
ScanReportemitted from scan planning andCommitReportemitted from commits, with environment context in commit reports - Reporters: A REST metrics reporter and an OpenTelemetry reporter, documented in a metrics reporting guide
Row-Level Operations and Maintenance🔗
- Merge-on-read writes: A
PositionDeltaWriterfor insert/reinsert splits,RowDelta.RemoveDeletesfor deletion-vector supersession, and aCommitEqualityDeletesconvenience API - Dynamic overwrite: A partition-match predicate for dynamic overwrite
- Compaction: Partial rewrites committed per group, compaction forced on file-scoped delete dead space,
CollectDeadPositionDeletesfor leader-side expunge,RewriteFilesable to add delete files, and tunables bounding compaction pipeline memory - Rewrite manifests: A
RewriteManifestsmaintenance task and CLI command with opt-in clustering - Sort order on write: The table's default sort order is now enforced on write
Transactions and Schema Evolution🔗
- Union schema evolution: Evolve a schema to the union of the current and an incoming schema by name, matching Java's
unionByNameWith - New transaction APIs:
Transaction.RemoveProperties,Transaction.ReplaceSortOrder,Transaction.AssertRefSnapshotID, and an explicit default spec/sort-order commit fence viaTransaction.AssertDefaultShape - Partition specs: Added data files can target any registered partition spec
- Branch-aware writes: Write planning resolves snapshots from the target branch head,
RollbackToSnapshotis scoped to the transaction's target branch, and multi-table commits are fenced with the target branch requirement - Atomicity: Transaction apply is now atomic and transactions are only marked committed after a successful commit
Catalog Improvements🔗
- REST endpoint negotiation: The REST client honors the endpoints advertised by
/v1/config - REST features: Rename view, custom namespace separators,
snapshot-loading-modesupport inloadTable, catalog labels on the load table/view read path, andoauth2-server-urifor the OAuth token endpoint - SQL UDFs: A SQL UDF metadata model package and read support through the REST catalog
- Purge table: A
PurgeTableoptional interface for physically deleting table files - Hadoop catalog: Moved onto the generic IO interface with support for blob filesystems and opt-in arbitrary schemes
- Hive catalog: Table drops and renames are locked, HMS table metadata is synchronized on commit, and lock retry behavior aligned with Java
- Conformance tests: A shared catalog conformance test suite now runs against every catalog implementation
Encryption Foundations🔗
Groundwork for Iceberg table encryption was added:
- Interfaces:
EncryptionManagerandKeyManagementClientinterfaces with an in-memory KMS - KMS registry: A KMS catalog-property registry selected through
encryption.kms-type/encryption.kms-impl
IO Improvements🔗
- Per-cloud packages:
io/gocloudwas split into per-cloud backend packages so applications only link the cloud SDKs they use - Azure:
ManagedIdentityCredentialsupport - S3: Connect timeout property, compatibility with GCS HMAC endpoints, and explicit
s3.*credentials overriding the context AWS config - GCS: The GCS client authenticates with supplied credentials
- Glue:
aws-profilesupport for the Glue catalog in the CLI config
Performance🔗
More than 120 performance PRs landed in this release, touching nearly every part of the read, write, and planning paths. Some of the highlights:
- Parallel reads: Large Parquet files are split into scan tasks by row-group byte range, controlled by
read.split.target-size - Scan planning: Snapshot manifests cached across scans, compiled filter and pruning plans cached, scan columns projected during manifest reads, and row limits honored during local planning
- Delete handling: Equality deletes indexed by partition and sequence, positional deletes indexed by path and partition, equality and position deletes loaded lazily per scan task, and equality deletes pruned by data-file metrics
- Row-group pruning: Parquet row groups pruned by dictionary
- Maintenance: Shared manifest reads deduplicated during snapshot expiration and orphan cleanup
- Writes: Parquet dictionary encoding enabled by default, matching Java Iceberg
Arrow Go v18.8🔗
iceberg-go now depends on Apache Arrow Go v18.8.0, up from v18.6.0 in 0.6.0 (#1504, #2002), so applications importing both modules will resolve arrow-go to at least v18.8.0. The new compute.VariantGet kernel backs variant extract, with a zero-copy fast path for fields shredded to the queried type. Variant columns are now written with the canonical arrow.parquet.variant Arrow extension name; 0.7.0 reads both the canonical and legacy parquet.variant names, but older iceberg-go readers will not recognize the new name during a mixed-version rolling upgrade.
Bug Fixes🔗
Nearly 400 bug fixes landed in this release, many of them hardening validation, JSON decoding, and defensive copying across the public API. Notable fixes include:
- Fix deletion vectors orphaned by data-file rewrites
- Fix OCC rewrite resurrection
- Fix
_row_idrenumbering when rows are dropped in a scan - Keep generated position deletes correct under row-group pruning
- Apply positional deletes with batch-local
Takeindices - Fix manifest pruning when partition source columns are dropped
- Fix incorrect pruning for
NOT STARTS WITHover truncated string partitions - Truncate strings by Unicode code points instead of UTF-8 bytes
- Retain ref head snapshots in expire snapshots
- Respect
gc.enabledduring snapshot expiration - Detect concurrently-removed data files in row-level deletes
- Fix handling of partition specs with nested fields
- Encode day-transform partition fields with the Avro date logical type
- Fix a big-endian alignment crash
- Fix js/wasm builds by dropping
ptermfrom the library path
Breaking Changes🔗
Some of these changes are breaking changes that need to be called out:
- Scan concurrency option (#1462): the misspelled
WitMaxConcurrencyscan option was renamed toWithScanMaxConcurrency. - Transform interface (#1252):
Transformgained a type-awareToHumanStrTypemethod, which externalTransformimplementations must add.ToHumanStr(any)is deprecated. - Hadoop catalog IO interfaces (#1185): the
MkdirIOandReadDirIOinterfaces were removed in favor ofMkdirAllandWalkDir. - Hadoop catalog cloud schemes (#1961):
catalog/hadoopno longer registers cloud schemes as a side effect. Users on cloud storage must blank-import the matchingiobackend package. The existingio/gocloudsymbols remain available as deprecated aliases. - Orphan cleanup (#1286): orphan file cleanup now requires the table's
FileIOto implementListableIO. - Parquet dictionary encoding (#931): data files are now written with dictionary encoding enabled by default.
New Contributors🔗
Welcome to all 33 first-time contributors: @mithilgirish, @paveon, @sungwy, @Revanth14, @rraulinio, @artbcf, @gerinsp, @jeremybarner, @strophy, @neu-hsc, @ebyhr, @abnobdoss, @AhmedDlshad007, @nanookclaw, @prafulla-kiran, @rasta-rocket, @the-onewho-knocks, @iremcaginyurtturk, @hankivstmb, @jx2lee, @1991santhu, @mattfaltyn, @dgvj-work, @dimakuz, @ryanworl, @technicolorbeat, @codeAnqiang-ma, @deepak-2605, @smaheshwar-pltr, @ajiteshsingh, @sgg, @KranzL, @shoemoney
Getting Involved🔗
The iceberg-go project welcomes contributions. We use GitHub issues for tracking work and the Apache Iceberg Community Slack for discussions.
The easiest way to get started is to:
- Try
iceberg-gowith your workloads and report any issues you encounter - Review the contributor guide
- Look for good first issues
For more information, visit the iceberg-go repository.