AWS Unveils Glue 6.0: A Major Leap Forward with 30% Price Reductions, Apache Spark 4.1, and Full Apache Iceberg v3 Support

SEATTLE — In a move set to reshape the economics and technical capabilities of cloud-based data engineering, Amazon Web Services (AWS) has officially announced the general availability of AWS Glue 6.0. The latest iteration of the fully serverless data integration and ETL (Extract, Transform, Load) service introduces a modernized runtime stack, deeper support for open table formats, and a significant 30% price reduction compared to previous versions.

Built on top of cutting-edge open-source frameworks—specifically Apache Spark 4.1, Python 3.13, and Scala 2.13—AWS Glue 6.0 aims to tackle the growing complexity of handling massive, semi-structured datasets while dramatically lowering the total cost of ownership for data-driven enterprises.


Main Facts: What is New in AWS Glue 6.0?

AWS Glue 6.0 represents one of the most comprehensive upgrades to the serverless data platform in recent years. Designed to address the modern demands of data lakes, data meshes, and real-time streaming architectures, the new release is anchored by three primary pillars: dramatic cost savings, a modernized execution engine, and complete Apache Iceberg v3 integration.

1. A 30% Reduction in Pricing

Perhaps the most immediate commercial draw for enterprise data teams is the structural price cut. AWS Glue 6.0 delivers a 30% reduction in pricing relative to earlier AWS Glue versions. Because data pipelines often process petabytes of information on a daily or hourly basis, this cost optimization directly improves operational margins for organizations scaling their data operations in the cloud.

2. Modernized Runtime: Spark 4.1, Python 3.13, and Scala 2.13

Under the hood, AWS Glue 6.0 has been re-engineered for speed and efficiency. By upgrading its core engine to Apache Spark 4.1, the service unlocks the latest performance optimizations and query execution enhancements developed by the open-source community. Furthermore, the inclusion of Python 3.13 and Scala 2.13 provides developers with modern language features, improved security patches, and faster execution times for custom ETL scripts and PySpark workloads.

3. Complete Apache Iceberg v3 Implementation

AWS Glue 6.0 provides the industry’s most complete implementation of Apache Iceberg v3 (built on Iceberg 1.11.0) on any fully serverless managed Spark service. The standout feature of this integration is the introduction of the native VARIANT data type complete with shredding support.

Traditionally, handling semi-structured formats like JSON, application logs, and event streams required cumbersome workarounds: storing data as strings, writing custom parsing code, or rigidly flattening schemas. When upstream schemas inevitably changed, these pipelines frequently broke, requiring manual intervention and creating duplicate copies of data.

With the new VARIANT data type and shredding capabilities in Iceberg v3:

AWS Glue 6.0 now available with 30% lower price and full Apache Iceberg v3 support | Amazon Web Services
  • No Schema Flattening Required: Engineers can ingest raw JSON and event data natively.
  • Elimination of Data Duplication: Organizations no longer need to maintain multiple parsed variations of the same dataset.
  • Accelerated Query Performance: Shredded VARIANT columns allow query engines to read specific attributes directly, achieving significantly faster read performance than traditional string-based semi-structured columns.
  • Resilience to Change: Upstream schema evolutions no longer fracture downstream data pipelines.

Chronology: The Evolution Leading to Glue 6.0

To understand the significance of AWS Glue 6.0, it is helpful to look back at how AWS has systematically evolved its serverless analytics and data integration portfolio to meet the demands of open-source table formats.

  • The Early Era of Managed ETL: AWS Glue was initially launched to provide a serverless alternative to managing dedicated Apache Spark and Hadoop clusters. It abstracted infrastructure management, allowing data engineers to focus purely on code and data transformations.
  • The Rise of Data Lakes and Open Formats: As data lakes migrated from flat CSV and Parquet files to transactional table formats like Apache Hudi, Delta Lake, and Apache Iceberg, AWS steadily introduced metadata management tools, such as the AWS Glue Data Catalog, and native integrations to support ACID transactions.
  • Incremental Runtime Upgrades: Over the subsequent years, AWS rolled out successive versions of Glue, progressively updating its Spark and Python runtimes to keep pace with the open-source ecosystem.
  • The Shift Toward Cost-Efficiency and Interoperability: Facing macroeconomic pressures and an industry-wide focus on cloud cost optimization, enterprise customers increasingly sought ways to do more with less. Concurrently, Apache Iceberg emerged as the de facto standard for open table formats.
  • August 2026 – The Release of AWS Glue 6.0: Culminating years of runtime optimization and architectural refinement, AWS releases Glue 6.0. It bridges the gap between ultra-low-latency real-time streaming, advanced semi-structured data management via Iceberg v3, and aggressive price points.

Supporting Data and Technical Architecture

The technical architecture of AWS Glue 6.0 is optimized to eliminate bottlenecks across storage, compute, and metadata cataloging.

Performance and Latency

The combination of Spark 4.1 and Iceberg v3 support allows AWS Glue 6.0 to handle demanding streaming and batch workloads. The platform now enables real-time data streaming operations with single-digit millisecond latency, bringing transactional data lake operations closer to real-time operational database performance.

Pricing Structure Breakdown

AWS Glue 6.0 maintains a transparent, pay-as-you-go economic model while lowering the base cost of execution:

  • ETL Jobs and Crawlers: Customers pay an hourly rate, billed by the second, for the compute resources consumed while discovering data (crawlers) or executing data transformation and loading scripts (ETL jobs).
  • AWS Glue Data Catalog: Metadata storage and access follow a simplified monthly fee model. To encourage adoption, AWS provides a generous free tier: the first one million objects stored and the first one million accesses are completely free.

Regional Availability

At launch, AWS Glue 6.0 is generally available today across all AWS Regions where AWS Glue operates. Developers can check specific regional capability rollouts via the AWS Builder capabilities tool.


Official Perspectives and Expert Commentary

Industry analysts and AWS community leaders have widely praised the release for addressing two of the most persistent pain points in modern data engineering: escalating cloud infrastructure bills and the operational overhead of managing semi-structured data schemas.

Cloud architects note that the introduction of Iceberg v3’s VARIANT shredding addresses a critical friction point for teams working with modern event-driven applications and microservices. Because microservices frequently emit JSON payloads with dynamic attributes, data engineering teams previously spent countless engineering hours maintaining brittle schema evolution scripts.

Furthermore, community voices have emphasized the frictionless nature of the upgrade path. Channy Yun, principal developer advocate at AWS, highlighted in the official release announcement that moving to the new version does not require massive code rewrites or complex API overhauls.

AWS Glue 6.0 now available with 30% lower price and full Apache Iceberg v3 support | Amazon Web Services

Implications: What AWS Glue 6.0 Means for the Data Engineering Industry

The launch of AWS Glue 6.0 carries profound implications for data teams, enterprise budgeting, and the broader data tooling ecosystem.

1. Democratizing Advanced Open Table Formats

By baking complete Apache Iceberg v3 support directly into a fully serverless managed service, AWS is lowering the barrier to entry for modern data lakehouses. Organizations that previously lacked the specialized engineering talent required to manually configure, tune, and maintain Iceberg tables can now leverage enterprise-grade capabilities out of the box.

2. Shift Toward FinOps and Cost-Conscious Architecture

With a 30% price reduction, AWS is aggressively positioning Glue to compete against alternative cloud data warehouses and third-party ETL platforms. Data leaders are increasingly tasked with demonstrating strict ROI on cloud expenditures. Glue 6.0 allows organizations to maintain high-performance, large-scale Spark workloads without proportionally increasing their cloud budgets.

3. Simplified Migration Paths

AWS has taken proactive steps to ensure that upgrading does not become an administrative nightmare. Organizations can adopt Glue 6.0 with minimal friction:

  • No API Changes Required: Existing scripts and orchestration tools continue to function without modification.
  • Flexible Invocation: Developers can target the new version using the standard --glue-version parameter via the AWS CLI, SDKs, AWS Glue Studio, or Amazon SageMaker Unified Studio.
  • Automated Upgrade Agents: For legacy jobs, teams can utilize the Spark upgrade agent within AWS Glue Studio or configure the auto-upgrade feature to seamlessly transition existing pipelines to Glue 6.0.

4. Integration with AI and Assistant Tools

Reflecting the modern developer experience, AWS has also ensured that Glue 6.0 is fully supported by modern AI-assisted engineering workflows. Developers can search documentation, verify regional availability, query APIs, and troubleshoot migration issues using the AWS MCP Server and associated plugins within their preferred AI-powered IDEs.


Getting Started

Data engineers and platform administrators can begin experimenting with AWS Glue 6.0 immediately.

To create a new job or upgrade an existing pipeline in the AWS Glue Studio console, navigate to the Job Details tab and select the version labeled Glue 6.0 – Supports Spark 4.1, Scala 2, Python 3. For interactive development, data scientists and engineers working within Jupyter notebooks or AWS Glue Studio notebooks can initialize the environment by setting 6.0 in the %glue_version magic command.

As data lakes continue to grow in volume, velocity, and variety, AWS Glue 6.0 positions itself as a robust, cost-effective cornerstone for the next generation of cloud-native data architectures.