Picture a cloud architect at a large company. Finance has just handed him a familiar assignment: next year’s budget needs credible cloud numbers, broken down by application and team, with a forecast leadership can stand behind.
He has reports. The reports show tag coverage against the company’s tagging standard, and a cost dashboard where 100 percent of the bill is allocated across a tidy list of applications. On paper, the assignment looks easy.
The problem is that he doesn’t believe the reports. Not because anyone is lying, but because he knows how the numbers are produced. Coverage means a tag exists, not that it’s right. The 100 percent allocation is achieved with catch-all buckets and evenly spread shared costs. When someone asks which team drove last month’s increase, or what one application actually costs end to end, the tidy dashboard has no answer. Network traffic all rolls up to something, because there’s no way to differentiate it.
He has no confidence in the numbers. And, more uncomfortably, he has no clear way to get confidence. That second problem turns out to be the interesting one.
What Staying in the Dark Actually Costs
Before the research, it’s worth being honest about the stakes, because “our tags are messy” undersells them.
The money is real. Waste survives where attribution is vague. Idle capacity, orphaned resources, and oversized instances persist precisely because nobody owns them on paper. Industry research consistently puts cloud waste at 20 to 30 percent of spend, and every month in the dark compounds.
Bad numbers drive bad decisions. Showback shapes behavior. When shared costs are spread evenly, efficient teams subsidize inefficient ones, and every downstream decision built on those numbers, unit economics, pricing, build versus buy, inherits the error.
Untagged means unowned, and unowned is a risk. A resource with no owner is a resource nobody patches, nobody backs up, and nobody can speak for during an incident. If sensitivity tags are missing, you can’t answer where regulated data lives. The tagging gap and the governance gap are the same gap.
The gap is widening on its own. New AI and marketplace consumption is landing in cloud bills faster than any tagging standard can be revised to describe it. Waiting doesn’t hold the problem steady; it grows it.
And there’s the room. The first time finance finds a catch-all bucket behind a “100 percent allocated” report, every other number you bring loses credibility. That trust is expensive to rebuild.
The Research Phase: Every Path Leads Back to Tags
So he does what good architects do. He researches how other organizations assess what they have, and looks for a better way to organize it. Three familiar options come up.
Option one: enforce tags in Terraform. Write or adopt a module that requires the standard tags at provision time, and non-compliant infrastructure can’t deploy. It’s a reasonable guardrail, and he nearly starts building it. But walking through it honestly, it governs the future, not the present. Everything that already exists stays exactly as untrusted as it was. Anything created outside the pipeline escapes it. And it can enforce that a tag is present, but not that it’s true. The moment a team reorganizes or a project ends, tags that passed validation quietly become wrong. He’d be adding maintenance burden to gain enforcement of a standard he can’t verify against reality.
Option two: deploy a FinOps tool. Apptio, Cloudability, Flexera, and their peers are capable products, and one of them is probably in his future regardless. But reading the implementation guides reveals an awkward dependency: these tools largely attribute costs through tags. Feed them a low-confidence tag foundation and they produce low-confidence output, formatted beautifully. Which means deploying one properly starts with a tag remediation project, and now the tool evaluation has become the very multi-quarter program he was trying to avoid.
Option three: audit it manually. Scripts, spreadsheets, interviews, maybe consultants. This is how it’s traditionally done, and it’s how his own FinOps reporting works today. He also knows the arc: the inventory takes a quarter or more to build, engineers leave and teams reorganize while it’s underway, and the result starts decaying the day it’s finished. Organizations that go this route tend to do it more than once, which tells you what the snapshot is worth.
Sitting with all three options, something dawns on him. Every path leads back to tags. And every problem with every path is really a problem with what tags are.
Questioning the Tag Itself
A tag is a piece of metadata a human was supposed to attach, describing something the cloud already knows.
Who created this resource? The cloud has that in its audit trail. What deployed it? The pipeline knows. What does it connect to, what depends on it, which accounts and projects does it live in, how is it actually used? All of that is observable. The tag is a hand-maintained sticky note summarizing evidence that exists, in full, one layer down.
Seen that way, the old way’s whole cost structure makes sense. Standards documents, enforcement modules, compliance reports, and reconciliation meetings are all machinery for getting humans to manually transcribe reality into metadata, and for detecting when the transcription goes stale. The machinery is expensive because the transcription never stops and the drift never stops.
For decades that machinery was necessary, because nothing could read the underlying evidence at scale. That’s the assumption that no longer holds.
The New Way: Read the Evidence, Then Fix the Tags
AI agents are very good at exactly the work that made visibility expensive.
An agent with read access to a cloud environment can build a live map of it: every resource, how everything connects, what’s load-bearing and what’s orphaned. Against that map it can baseline tag coverage in days rather than quarters, and not as a snapshot but as a continuously current view. Where tags are missing or suspect, it can propose values by reading the evidence: creator, deployer, connections, usage. Where attribution is vague, it can show exactly which spend is genuinely allocated and which is hiding in a catch-all.
That map has a name: topology. And a live, up-to-date topology is quietly becoming one of the most critical backbones of a strong cloud posture. It’s what visibility rests on, what makes AI safe to use against your environment, and what everything else in this article depends on. Tags were always a stand-in for it, a version everyone agreed not to fully trust, not least because tags slip out of sync quickly and are hard to bring back. That arrangement can’t hold. As more of the work moves to AI, an airtight topology that stays current stops being a nice-to-have and becomes the requirement.
An agent is only as safe as its guardrails and as accurate as its context, which is why the topology comes first and every change waits for approval.
This is the “turn on the light” moment. Not a migration, not a new platform to roll out, not a change to how teams work. A lens applied to the cloud you already have, using the credentials you already have. First you see clearly. Then, and only then, you organize, clean out, and make changes, with the confidence that comes from knowing what’s actually there.
The sequencing matters. Cleanup without visibility is how zombie resources survive every audit, and how a well-meaning tag remediation breaks a production dependency nobody knew existed.
Don’t Change Your Tools. Change What They Can See.
Notice what this does to the three options from the research phase. None of them were wrong. They were just out of order.
The Terraform guardrail becomes worth building once there’s a verified baseline to enforce against, instead of a standard describing an imaginary cloud. The FinOps tool becomes worth deploying once the tag foundation feeding it is trustworthy, and its findings get more useful still when an agent like Oscar ingests them, applies context from the live map, and turns them into a prioritized, ROI-ranked task list: what to fix, why it’s safe, what it’s worth, and who owns the decision. Even the manual audit’s goal, a complete inventory, arrives as a byproduct, except it stays current instead of decaying.
Nothing gets replaced. The tools you’ve already invested in keep doing what they’re good at, and the work flows through the systems your teams already live in, like Jira and ServiceNow, rather than yet another queue.
What Finance Gets Out of It
Back to the original assignment: credible budget numbers. The journey from here is incremental, and each step pays for the next.
- Turn on the light. Connect read-only, build the live topology, and baseline reality: tag coverage, ownership gaps, attribution quality, and where the spend actually goes.
- Set the standard against reality. The tagging standard stops being a document describing an imaginary cloud and becomes a measurable delta: here’s the gap, here’s the plan to close it, here’s this month’s progress.
- Organize and clean out. Assign the inferred owners. Remediate the highest-value gaps first. Retire what nothing depends on. Every change proposed with evidence, executed with approval.
- Keep the light on. The step the old way never reached, because maintaining visibility manually cost as much as building it. A live map doesn’t drift. New resources show up already evaluated, and enforcement tooling becomes a guardrail on a road you can actually see.
For finance, that cadence changes everything. Cost visibility stops being a promised future state at the end of an eighteen-month program and becomes a monthly, measurable improvement, with a budget built on attribution someone can actually defend in the room.
Advances in technology don’t just make old approaches faster. They change which approaches make sense at all. Transcribing your cloud into metadata by hand, and policing the transcription forever, made sense when nothing could read the evidence directly. It doesn’t anymore.
Where to Start: The First 30 Days
None of this requires a program office. A realistic first month looks like this.
Week 1: Write down the questions you can’t answer. Which team drove last month’s increase? What does application X cost end to end? Who owns the ten most expensive untagged resources? That list is your success test. Gather the tagging standard and current reports alongside it.
Weeks 1 to 2: Baseline with evidence, not surveys. Give an AI agent read-only access and let it map the topology and score your actual tag coverage and attribution quality against the standard. Modern agents need no tagging to build the map, and nothing changes in your environment. No new platform, no migration, credentials you already have.
Weeks 2 to 3: Publish the honest gap. Real coverage number, the top attribution gaps ranked by dollars, unowned spend. Take it to finance. Counterintuitively, this builds credibility rather than spending it: an honest 60 with a plan beats a hollow 100.
Weeks 3 to 4: Close the highest-value gaps first. Proposed owners and tags backed by evidence, every change approval-gated, and each fix verified in the next report. The loop isn’t done when a recommendation is sent. It’s done when the number moves.
Day 30 and beyond: Keep the light on. Turn the baseline into a continuous one, add provision-time enforcement (yes, that Terraform module, now aimed at a verified standard), and give finance a monthly delta instead of an annual apology.
You can run this plan with scripts and patience; teams have. The reason it now fits in a month instead of eighteen is that reading the evidence is the part AI agents do well. That’s the part Oscar was built for.
Turn On the Light in Your Cloud
Oscar Ops is an AI cloud engineering agent that runs locally on your workstation, connects to your cloud with the credentials you already have, builds a live Cloud Intelligence Graph of your environment, and proposes actions that execute only with your approval.
And soon, Oscar will operate inside your LLM of choice, whether that’s Claude, Claude Code, Codex, or Cursor. Connect Oscar to the AI you already use, and it brings its guardrails and live topology along, so this work happens right where you already work. Safety and accuracy are the two concerns that stop most teams from pointing AI at their cloud, and rightly so; with an airtight topology underneath and every change approval-gated, they stop being deal-killers.
If your organization is staring at a budget request, a messy bill, or a plan to script your way to compliance, start by seeing what you actually have. Meet Oscar to learn how it works, or try Oscar Ops and baseline your own environment.
Key Takeaways
Key findings
- ✓A cost report showing 100 percent allocation isn't the same as accurate attribution. Catch-all buckets hide the fallout instead of explaining it.
- ✓Tag enforcement tools only govern what gets created going forward. The existing estate, the drift, and the ownership gaps stay dark.
- ✓Most FinOps tools attribute costs through tags, so deploying one on a low-confidence tag foundation turns into a tagging project first.
- ✓AI agents changed the economics: a baseline of your topology, tag coverage, and attribution gaps now takes days instead of a multi-quarter program.
- ✓Publishing the honest gap builds credibility with finance instead of spending it. An honest 60 with a plan beats a hollow 100.
- ✗If your budget numbers depend on humans remembering to follow a tagging standard, those numbers are already drifting.