mayukh.dev /work/bulk-version-update.agent

Bulk Terraform Version Update using Agentic Workforce

33%80%+
of HCP workspaces on the latest Terraform version, within two months of release
Role
Lead UX designer
Team
2 UX designers, 2 developers, 2 Product Managers
Timeline
Jan – April 2026
Company
IBM HashiCorp

Organizations running Terraform at scale can end up with hundreds of workspaces, each pinned to a different mix of core, provider, and module versions. As security advisories and deprecations stack up, platform teams struggle to keep up with tracking and remediation across every workspace.

I worked with the platform engineering team to understand how they currently managed versions and updates, where the process broke down, and what information they needed to prioritize and act quickly. My responsibilities included mapping the existing workflow, identifying the biggest bottlenecks, and shaping a clearer, more scalable approach to monitoring and remediation.

When platform engineers need to align workspaces with a new Terraform version, they want to safely validate and execute upgrades at scale so that they can focus on core infrastructure management and avoid the operational tax of manual updates or the risk of production downtime.
Verify at Scale: As a Workspace Operator, I want to validate the impact of a new Terraform version across my entire environment to achieve the confidence needed to upgrade without the fear of breaking infrastructure.
Fleet-Wide Alignment: As a Workspace Operator, I want to bring all my workspaces up to a target version in a single workflow to achieve a standardized environment without the manual toil of per-workspace configuration.
Improve adoption of Terraform flagship capabilities like Search and Actions which are only available in the latest versions.

Impact of Bulk version update feature

Increase in number of workspaces with latest Terraform version

Following the release of this feature, the proportion of HCP workspaces running the latest Terraform version increased from 33% to more than 80% within two months. This rapid adoption demonstrates strong customer uptake and validates the effectiveness of the new upgrade experience.

Increase in RUM (Resources Under Management) – increased Business Revenue

The feature release led to a marked increase in Resources Under Management (RUM). Because Terraform’s revenue scales with the number of managed resources, this RUM growth translated into a clear increase in business revenue.

Built with Claude Code and the Helios MCP server

Demo developed using Claude Code and Helios Designs System MCP server. The demo is best viewed in wide desktop screens. Not suitable for mobile screens.

Live product demo link

“Upgrading modules and providers becomes hell week. The full team ends up spending all their time upgrading and checking if anything has broken or not.”

John Weignar · JP Morgan & Chase Bank (Solution Architect)

Why upgrading is a complex engineering exercise

Row of falling dominoes
STEP 01
Core upgrade
Requires a certain version of a provider.
STEP 02
Provider upgrade
May introduce breaking changes to resource arguments.
STEP 03
Module update
The module author must release a new version to fix those breaking changes.
STEP 04
Workspace update
Operators repoint to the new module version, often triggering a resource replace instead of an update.
Dependency web
A single shared networking module upgrade can force dozens of downstream application workspaces to upgrade simultaneously to avoid configuration drift or plan failures.
The state trap
Versioning isn't just about the code, it's about the state. Once a state file is touched by a newer version of Terraform Core, older versions often cannot read it again, making rollbacks difficult.

HashiCorp already provides tools to upgrade workspaces between versions, and has previously shared scripts that address parts of the process. With an AI agentic workforce, we can now offer a unified experience that upgrades workspaces fully autonomously with minimal human intervention.

Personas of the Terraform ecosystem

Who is affected by the bulk update?

Architects
Platform Engineers

Define the upgrade strategy, manage the CI/CD pipelines (e.g., Atlantis, Terraform Cloud/Enterprise), and ensure the Core version stays secure across the fleet.

Librarians
Module Authors

Maintain the internal registry. When a provider changes (e.g., AWS 4.x to 5.x), they must update the building blocks so teams can consume them safely.

Translators
Provider Authors

Usually cloud vendors or community maintainers who release the API updates. Their release cycle is the upstream trigger for all other personas.

Consumers
Workspace Operators

App developers or SREs managing specific environments (Dev/Prod). They feel the brunt of the domino effect when they refactor workspace code to match new module requirements.

John upgrades hundreds of workspaces

User journey: four phases with actions, mindset quotes and an emotion curve from motivated to overwhelmed to relieved
Make the Workflow Topology Visible
Map the “Team”: Use a visual pipeline, DAG (node graph), or Kanban board so users see which agent is doing what.
Separate Process from Artifact: Keep messy agent back-and-forth/logs on one side (or in the background) and the final deliverable/canvas on the other.
Use Progressive Disclosure for Observability
Default: High-level milestone (e.g., “Step 2/4: Fact-checking draft”).
One click: Readable summary of agent decisions/debates.
Deep dive: Raw tool calls, API responses, and logs for debugging.
Enable Surgical Error Recovery
Pinpoint the Failure: Visually highlight the exact agent/node that failed and explain why.
Edit & Resume: Never force a full restart. Let users fix the broken intermediate data and click “Rerun from this step.”
Design for Long Latency & Async Runs
Milestones over Spinners: Show progress via completed sub-tasks, not indeterminate loading spinners.
Async Handoffs: Allow users to navigate away and receive notifications (Slack, email, in-app alerts) when tasks finish or need input.
Early Previews: Stream partial artifacts as they are completed by individual agents.
UX RESEARCH · SSE
Core and provider upgrades
01 · Terraform Core upgrades & the scale problem
“Upgrading a single workspace in the UI is trivial. Upgrading hundreds across isolated business units is an operational nightmare.”
— Fiona MacLeod, Lead Cloud Platform Engineer, SSE
02 · Provider upgrades & the "work cliff"
“When a provider drops a breaking change, we hit a work cliff. Suddenly we're staring down dozens of broken modules with zero engineering capacity to refactor them.”
— Liam Gallagher, Senior DevOps & Automation Engineer, SSE
UX implications
Upgrade workflows should be feature-driven or deprecation-driven, not routine maintenance prompts.
Multi-select workspace actions, batch upgrade campaigns, and organizational rollout progress trackers.
Early-warning CLI alerts and dashboard notifications that flag deprecation notices before upstream providers release breaking versions.
UX RESEARCH · SSE
Modules, testing and AI guardrails
03 · Module lifecycle, deprecation & testing gaps
“A code diff tells me what syntax changed. It doesn't tell me if a module bump is going to secretly destroy and recreate a production resource.”
— John Kuldeep, Principal Solutions Architect, Cloud & Platform Services, SSE
04 · AI & automation guardrails
“AI is welcome to write our tests and scan for syntax compatibility, but autonomous deployment is an absolute hard line. A human eye on production is non-negotiable.”
— Liam Gallagher, Senior DevOps & Automation Engineer, SSE
UX implications
Progressive deprecation workflows (soft warnings in CI → scheduled plan-phase blocks → hard deprecation).
Plan-level behavioral comparison views showing expected resource lifecycle changes (create/update/destroy) between module releases.
AI features must be designed as copilots/advisors (drafting tests, explaining diffs) rather than autonomous execution agents.
UX RESEARCH · JPMORGAN CHASE
Module dependencies and custom workarounds
02 · Module dependency hell & "hell weeks"
“Our modules are tightly intertwined but updated at totally different speeds. We literally lose an entire week every quarter just untangling broken module dependency chains.”
— David Sterling, Lead Cloud Enablement Engineer, JPMorgan Chase
03 · Custom workarounds, internal AI & Sentinel gaps
“We run over 1,600 Sentinel policy checks in a single pipeline to ensure governance, yet native tooling can't even tell us if an engine upgrade will break our syntax.”
— Sarah Chen, Staff Platform Engineer – Cloud Governance & Compliance, JPMorgan Chase
UX implications
Dependency visualization trees displaying downstream consumers and cross-module version compatibility.
In-product state migration suggestions (auto-generating moved and removed blocks).
Native multi-engine and multi-provider compatibility matrix testing within the private registry.
Private/bring-your-own-model (BYO-LLM) integration supporting iterative automated plan repair and syntax remediation.
UX RESEARCH · JPMORGAN CHASE
Enterprise scale, trust and the 0.14 deadlock
01 · Enterprise scale, trust & the 0.14 deadlock
“Compatibility promises don't mean anything when you're managing 100,000 workspaces. With no incentive to move, we stayed on 0.14 until vendor deprecation forced our hand.”
— Marcus Vance, Executive Director – Enterprise Cloud Platform, JPMorgan Chase
UX implications
Deprecation timelines and EOL milestones tied directly to enterprise compliance status dashboards.
Automated syntax refactoring guidance identifying deprecated experimental syntax.
"Strong Validation" batch testing: running speculative plans at scale in a non-destructive sandbox mode.

Impact of Terraform Bulk Version Update

The Terraform Bulk Version Update feature shows how an agentic workflow can be designed to earn trust from infrastructure teams. By pairing automation with clear guardrails—scoped permissions, auditable change logs, and human-in-the-loop checkpoints—the experience reduces manual upgrade toil while keeping teams in control. The result is faster, safer adoption of the latest Terraform versions across hundreds of workspaces, turning a risky, repetitive operation into a routine, low-friction task. For UX, this underscores a core principle for enterprise tooling: agentic systems must be as transparent and verifiable as they are powerful, so teams can confidently let automation handle scale.

Next: AI Memory Profiler Back to portfolio