When AI meets
sustainability.

Since you opened this page, AI data centers have consumed
0 kWh
Based on IEA 2025 estimate: ~945 TWh/year by 2030  ·  ~30 MWh per second globally
A DataCamp report

Every company in the world is answerable to three things.
Most are only beginning to understand how AI changes all three.

E
Environmental
S
Social
G
Governance

Hover over each icon to reveal what's broken.

E, S, and G. For two decades, this three-part framework has been how investors, regulators, and the public ask a simple question: is this company doing right by the world? The answers become scores that shape capital flows worth trillions of dollars.

The framework was designed around annual disclosures, human decision-makers, and a relatively stable set of risks. It wasn’t built for AI systems that make millions of automated decisions, consume extraordinary energy and water, and reshape labor markets at a speed reporting cycles can’t track.

AI isn’t one thing — it’s at least three, each with different implications for how ESG frameworks need to evolve.

Type 01
Predictive AI

Learns from historical data to forecast, score, and classify. The oldest type — it's been running credit models and risk assessments for 20 years. It already powers most ESG ratings you've ever seen.

Already embedded in ESG scoring
Credit risk scoring — live model
Credit history
—
Employment
—
Location
—
Account age
—
Risk score
—
Model recalculates automatically. No human in the loop.
Type 02
Generative AI

Produces content: text, code, images, analysis. The kind behind ChatGPT, Copilot, and Gemini. The one your team is already using. Also the one that consumes extraordinary energy and water to run at scale.

Changed the game in 2022
Generating ESG disclosure draft…
|
Energy used
0.00 Wh
Type 03
Agentic AI

Systems that don't just respond — they act. Make decisions. Execute tasks. Adapt without waiting for a human to press go. The newest arrival, and the one where the accountability questions haven't been answered yet.

The frontier, right now
Agent executing autonomously
📧
Scan emails
→
📝
Draft board memo
→
💸
Approve $2.3M payment
→
🏛
File SEC disclosure
Human approvals 0 / 0 steps
Each one is changing ESG — in two directions at once.

The same technology is doing both things at once.
Which side wins depends on what gets measured.

Making it harder
More energy to run.
More jobs exposed.
Harder to audit.
Making it better
More data to measure.
Faster to report.
Smarter to govern.

The companies that understand this
are measuring both sides — not just the one that looks good.

explore the three pillars ↓
01
Environmental
The energy story
nobody’s telling
01 · Environmental

Every time you type something into ChatGPT, it uses electricity. That probably doesn’t surprise you — everything uses electricity. But the amount matters more than most people realize. In June 2025, OpenAI’s CEO Sam Altman finally put a number on it.

ChatGPT alone handles about one billion queries every single day. That number is what makes the per-query figure significant.

Energy per AI interaction
Sam Altman sketch
“

The average query uses about 0.34 watt-hours, about what an oven would use in a little over one second, or a high-efficiency lightbulb would use in a couple of minutes.

Sam Altman CEO, OpenAI  ·  The Gentle Singularity, June 2025
At 1 billion queries a day, ChatGPT consumes roughly 340,000 kWh[1] — enough to power about 11,500 US homes. That’s text only. Relative to a standard Google search, a text AI query already uses 11× more energy — an image generation uses 130× more, and an agentic task up to 1,300× more.[2]
Data center buildout
Global electricity consumed by data centers, 2010–2030
Scroll through the story →
2018 — baseline

Data centers were a solved problem.

In 2018, global data centers used 200 TWh per year — about 1% of global electricity.[3] The curve had been flat since around 2010 — efficiency gains kept pace with demand despite workloads growing 550%. Nobody was worried.

2019 — 2022

The curve was already bending.

Before ChatGPT, the flatness was already ending. Cloud computing scaled, video streaming surged, and early ML workloads — recommendation engines, fraud detection, image recognition — quietly consumed growing energy. The combined electricity use of Amazon, Microsoft, Google, and Meta more than doubled between 2017 and 2021.[4] By 2022, global consumption had risen to an estimated 240–340 TWh — up to 70% above the 2018 baseline.

2022 — inflection

Generative AI arrived and the curve broke.

ChatGPT launched in November 2022. Within months, every hyperscaler was racing to deploy. The flat line bent sharply upward. Electricity demand jumped 17% in 2025 alone[5] — four times the rate of global electricity growth.

2024 — 2026

$250 billion in 2024. $350 billion in 2025.

Microsoft, Google, Amazon and Meta spent a combined ~$250 billion[6] on capital infrastructure in 2024 — not on models or products, but on the physical buildings, power systems, and cooling equipment required to run them. In 2025, that figure rose to over $350 billion.[7] For 2026, the four companies have guided a combined spend approaching $700 billion.

2026 — the constraint

Money isn’t the problem. Energy is.

The hyperscalers have the capital. The grid doesn’t have the power. In Ireland, data centers consumed 22% of the country’s electricity in 2024 — more than every home in the country combined — forcing a moratorium on new Dublin connections.[8] In Frankfurt, data centers account for up to 40% of city power demand, with all available capacity already allocated.[9] New projects are being delayed not by budgets, but by grid queues that can take years.

2030 — projection

945 TWh. The IEA’s base case.

More than double 2024 levels. More than Japan uses in a year. And this is the conservative scenario — assuming the current pace holds, not accelerates. The “lift-off” scenario puts it at 1,260 TWh.

Then there’s the water.

Energy isn’t the only resource data centers consume. They also need enormous amounts of water to stay cool — and much of that water evaporates and is gone, drawn from local supplies that are often already under stress.

Google’s data centers — 2024 water use[10]
1 Olympic
swimming pool
consumed every hour
~34 pools per day · 12,400+ per year
Each icon = 1 Olympic pool (660,000 gal)  ·  34 pools = one day  ·  scarce-watershed pools in red

To understand the scale: Google used 8.1 billion gallons[10] of water across its data centers in 2024 — a 32% jump from 2023, driven by AI cooling demand. Most of it evaporates in cooling towers and is gone — not returned to local supplies.

It’s not just operations. Training GPT-3 evaporated an estimated 700,000 liters of freshwater[11]. The same model consumes roughly one 500ml bottle of water per 10–50 responses, depending on where and when it runs.[11] That sounds small. At a billion queries a day, it adds up fast.

The building craze

The demand on paper is one thing. What’s being built in the desert is another.

AI’s energy story isn’t just about queries and TWh on a spreadsheet. It’s being built into the physical landscape right now, at a scale that has no real precedent.

Case in point — largest campus under construction
Switch Citadel Campus — Tahoe Reno, Nevada
Aerial view of Switch Citadel Campus, Tahoe Reno, Nevada
Switch Citadel Campus, Tahoe Reno, NV · PR Newswire / Switch
650 MW at full buildout[12]

7.2 million square feet on a 2,000-acre site next to Tesla’s Gigafactory. At full capacity, it would draw more power than the entire city of Tallahassee, Florida — whose all-time peak electricity demand is 633 MW[13] for a city of 200,000 people. Or think of it this way: roughly half the peak load of Reno, Nevada (population ~250,000) — from a single data center campus.

And this is just one campus. NV Energy reports tech companies have requested between 21,000 and 22,000 MW of new capacity across Nevada[14] alone — more than triple the state’s current grid.
The other side of the ledger

AI is genuinely being used to cut energy use and reduce emissions — mostly in industries outside the data center itself. Here’s what’s real, and what the caveats are.

33×
Google reduced energy per Gemini prompt 33-fold[15] in 12 months (May 2024 → May 2025). Efficiency improving faster than any prior energy technology.
Total queries are growing faster than efficiency gains. Net consumption is still rising.
26M t
CO₂ reductions enabled by Google AI products in 2024 — more than Google’s own 11.5M tonne footprint.[10] The biggest lever isn’t inside the data center. It’s supply chains: AI-optimised shipping routes, load consolidation, demand forecasting that cuts overproduction, and smarter grid dispatch. Maersk and DHL have published significant fuel savings from route optimisation alone.[16]
Downstream user reductions, not Google’s own. Self-reported, not independently verified. Supply chain gains are real but depend heavily on adoption depth.
3–10 pp
IEA estimate of the reduction in energy's share of production costs (pp = percentage points)[17] in energy-intensive industries through proven AI applications. Real gains, real facilities.
Most energy sectors are not yet deploying AI at scale. Digital skills and data are the main barriers.
“
The IEA was early in recognizing that there is no AI without energy — and that countries that provide secure, affordable and rapid access to electricity will be one step ahead.
— Fatih Birol, IEA Executive Director — Key Questions on Energy and AI press release, April 2026
How the 33× actually happened

The techniques making AI inference cheaper — and their limits.

Efficiency gains don’t happen by accident. They come from a stack of techniques applied across hardware, model architecture, and how inference is actually served. Each has a different ceiling.

Smaller & sparser models

Quantization — compress weights from 32-bit to 4-bit. A 100GB model becomes 25GB. Faster, cheaper, modest accuracy loss.

Mixture of Experts (MoE) — only a fraction of the network activates per query. Same output, fraction of compute. GPT-4 reportedly uses this; Mistral’s Mixtral confirms it.

Distillation — a large model trains a smaller one to mimic it. DeepSeek R1’s distilled variants match much larger models at a fraction of the cost.[18]

Ceiling: accuracy degrades on complex tasks.

Dedicated silicon

Groq & Cerebras — inference-only chips with no memory bottleneck. Sub-millisecond latency, lower power than GPUs.[19]

AWS Trainium & Inferentia — Amazon’s in-house chips for training and serving. Meaningfully cheaper per token than H100 instances for steady workloads.[20]

Google TPUs — Gemini inference runs on TPUs, not GPUs.[21] One reason Google’s per-query efficiency numbers look better than peers.

Ceiling: most of the world still rents Nvidia H100s by the hour.

Smarter inference

KV-cache — attention state from previous tokens is stored, not recomputed. Why cached prompts are cheaper in most APIs.

Speculative decoding — a small draft model generates candidates; the large model only verifies. 2–3× throughput, identical output quality.[22]

Continuous batching — groups multiple requests into one GPU pass, filling gaps dynamically. Major throughput gain at scale (vLLM pioneered this).

Ceiling: infrastructure-level — app developers get these gains automatically, or don’t, depending on the provider.

The Jevons paradox problem

Every one of these techniques is real and measurable. The problem is that cheaper inference means more inference. When a query costs 33× less, organisations run 33× more queries — or more. Efficiency gains that don’t reduce absolute consumption are productivity gains, not sustainability gains. The IEA’s 945 TWh projection already assumes continued efficiency improvement. It still sees demand nearly tripling.

AI is the largest new source of energy demand on the planet.
It’s also a powerful tool for reducing energy demand elsewhere.

The tension here is real and doesn’t resolve cleanly. AI is simultaneously driving up energy demand at an infrastructure level and enabling measurable reductions in how energy-intensive sectors operate. For companies reporting on climate, the honest accounting requires tracking both — the emissions attributable to AI workloads and the savings enabled by AI applications — rather than reporting only whichever number looks better.

02
Social
The people story
hiding in the infrastructure
02 · Social

When people talk about AI and society, they usually jump straight to jobs. But there’s a social cost that almost nobody is talking about yet: what AI infrastructure is doing to ordinary people’s electricity bills.

PJM grid capacity auction — mid-Atlantic US
2024–25
$2.2B
2025–26
$14.7B
Other grid costs Data center demand: $9.3B (63%)
Nearly 7× increase in one year. Data centers drove 63% of it. Source: PJM Inside Lines ($2.2B) / PJM Independent Market Monitor, via IEEFA ($14.7B)

What does that mean for ordinary people? Virginia's State Corporation Commission approved a $11.24/month increase[23] for 2026 — part of a broader rate case driven by infrastructure costs including, significantly, data center load growth. Nationally, a landmark April 2026 EIA federal report found U.S. utilities disconnected power to residential customers 13.4 million times[24] in 2024 — nearly four times higher than earlier partial-state estimates had suggested. And in areas near major data center clusters, wholesale electricity prices have increased 267%[25] between 2020 and 2025.

Ashburn, Virginia
Population: 46,000. Data center count: 200+.
Electricity use: more than most entire US states.
The town you’ve never heard of is one of the largest power consumers in the country — and its residents are now paying extra for it.
The past  /  2004–2022

Before ChatGPT, AI was already deciding your life chances.

The anxiety about AI and fairness didn’t start with ChatGPT. It started two decades earlier, with a quieter category of AI that most people never saw — predictive systems making high-stakes decisions about jobs, loans, bail, and benefits. In 2016, mathematician Cathy O’Neil named the problem in a bestselling book: Weapons of Math Destruction.[26] The models weren’t neutral — they were opinions, encoded in math, running at scale, with no right of appeal.

Falsely flagged high-risk Black 45% White 24% ↑ nearly 2× the false-positive rate
Recidivism scoring  /  2000s–present

COMPAS, used by US courts to predict re-offending, was found to falsely flag Black defendants as high-risk at nearly twice the rate of white defendants — while both groups had similar actual re-offending rates.[27] Judges used its scores in sentencing. Defendants couldn’t see the algorithm. Northpointe disputed this — “fairness” depends on which metric you measure.

Training data fed to the model Male 74% Female 26% ↓  Model learned: male = qualified
Hiring algorithms  /  2014–2017

Amazon’s resume screener taught itself that male candidates were preferable — penalising the word “women’s,” downgrading graduates of women’s colleges.[28] To their credit, Amazon identified it and scrapped it. Most companies never built that feedback loop. The bias wasn’t programmed — it was inherited.

Weekly care hours allocated Before 56 hrs After 32 hrs  −43% ↓ hours cut — no explanation given
Benefits & welfare  /  2010s

Arkansas deployed an algorithm to allocate home care hours for disabled and elderly residents. Hours were cut with no explanation. Residents sued. A federal court found the system violated due process. The state couldn’t explain how it worked.[29]

The self-reinforcing loop More patrols More arrests Prediction confirmed
Predictive policing  /  2011–2020s

PredPol trained on crime data from already over-policed neighbourhoods — creating a feedback loop: more patrols → more arrests → confirmed prediction → more patrols. Santa Cruz banned it in 2020. Los Angeles scrapped it the same year.[30] The pattern had already spread to dozens of cities.

“

Models are opinions embedded in mathematics.

— Cathy O’Neil, Weapons of Math Destruction, 2016

The pattern is consistent: a model trained on historically biased data, applied at scale, producing decisions that are hard to contest and harder to audit. These weren’t bugs — they were features, optimising for accuracy while distributional harm fell on people with the least power to push back.

The EU AI Act[31] — passed in 2024, with bans already active since February 2025 and full enforcement from August 2026 — prohibits social scoring and real-time biometric surveillance in almost all contexts (Article 5), and classifies hiring tools, credit scoring, and criminal risk assessment as high-risk, requiring mandatory transparency and human oversight. It took twenty years of documented harm to get there.

ESG frameworks are only now starting to ask the same questions the AI Act finally answered.

The present  /  2023–now

GenAI is now affecting actual jobs.
But not the way the headlines say.

~55K
US jobs explicitly attributed to AI in 2025, out of 1.21M total layoffs
That’s 4.5%.
20.4%
of Q1 2026 tech layoffs explicitly linked to AI — up from 8% in 2025
The share is accelerating.
55%
of companies that made AI-related cuts reportedly regret the decision

The reality is messier than either headline — "AI is destroying jobs" or "AI layoffs are a myth." There are at least three distinct things happening at once, and conflating them obscures what actually matters.

1  /  Cutting headcount to fund the infrastructure
META 2025: CAPEX UP, HEADCOUNT DOWN Capex ↑ $72B Jobs ↓ 3,600

Meta laid off 3,600 people in early 2025 while announcing $72B in AI capex the same year.[32] Google cut 12,000 jobs in 2023 citing pandemic overhiring — then spent $32B on data centers that year, $52B in 2024, and $91B in 2025.[33]

2  /  Restructuring on a belief about the next three years
THE CONVICTION GAP Evidence low Conviction high

Some companies are restructuring ahead of AI capability they expect to arrive — not capability that exists today. The layoffs are real; what’s driving them is conviction, not demonstrated performance.

HBR, Jan 2026:[34] “Companies are laying off because of AI’s potential — not its performance.”
3  /  Macro corrections with an AI label attached
OF 1.21M 2025 LAYOFFS — SHARE ATTRIBUTED TO AI ~55K AI 1.15M other reasons

Gartner found most 2025 layoffs were unrelated to AI — driven primarily by federal government actions and post-pandemic hiring corrections.[35] Only ~55K of 1.21M layoffs were explicitly attributed to AI.[36] “AI efficiency” is a better investor narrative than “we overhired.”

What makes this hard to track in an ESG context is that the three buckets are almost never disaggregated in company reporting. Workforce disclosures typically show headcount, not rationale. A company that cut 10% of its workforce to fund data center construction looks identical in the data to a company that laid off the same 10% because it overhired in 2021.

The structural signal — job postings over 18 months
BLS 10-year employment projections. Declines in both roles predate AI being widely cited as a cause. Source: U.S. Bureau of Labor Statistics, Occupational Outlook 2024–2034

“AI replaces tasks, not jobs.”
Mostly true. Dangerously wrong at the edges.

McKinsey estimates that 60% of jobs have at least 30% of their tasks that AI could automate.[37] For most workers — a manager, a nurse, an engineer — that means AI handles parts of the job faster, while the worker shifts toward higher-value tasks. Output rises without headcount falling. But that pattern breaks down for one category of jobs: roles where the task being automated is the entire job, with little left over.

Financial analyst
JPMorgan, BlackRock, Goldman
AI handles data gathering, screening, and report drafting → analyst focuses on judgment, client relationships, and edge cases. Output rises. Headcount stays roughly flat.
Task automation → job evolution
Customer service rep
Block, Salesforce
AI resolves 70–80% of queries. The remaining reps handle harder tickets — which take longer per case. Headcount falls, but not proportionally to query volume. The job shrinks significantly; it doesn't simply vanish.
Task automation → job elimination
Yuval Noah Harari sketch
“

As algorithms push humans out of the job market, wealth might become concentrated in the hands of the tiny elite that owns the all-powerful algorithms, creating unprecedented social and political inequality.

Yuval Noah Harari Author, 21 Lessons for the 21st Century, 2018
Written in 2018 — before generative AI. The scenario Harari described is now measurable in labour market data. The WEF projects 92 million jobs displaced by 2030[38], with gains concentrating in roles requiring advanced technical education that most displaced workers do not hold.
Named examples  /  2024–2026

The companies making that bet explicitly.

2024
Klarna
Deployed an AI chatbot it claimed could do the work of 700 customer service agents — reducing headcount through a hiring freeze and attrition rather than direct layoffs. Headcount fell from 5,000 to 3,800 during the same period.
Sept 2025
Salesforce
CEO Marc Benioff: customer support headcount cut from 9,000 to 5,000 via agentic AI. “There are more than 100 million leads that we have not called back at Salesforce in the last 26 years. But we now have an agentic sales [team] that is calling back every person that contacts us.”
Early 2026
Block
4,000 layoffs concentrated in customer support. AI resolves 70–80% of inquiries without human intervention. CEO Jack Dorsey explicit about AI capability as the cause.
March 2026
Atlassian
1,600 layoffs — 10% of the workforce. Cannon-Brookes framed the cuts as a structural move to become an AI-first company, redirecting resources toward AI development and enterprise sales.
2025–2028
Wall Street
Bloomberg: 200,000 banking jobs targeted over 3–5 years. Entry-level and back office — historically the first rung for recent graduates.
The other side

New roles are being created. Here’s what the numbers actually show — and what they leave out. The people losing jobs and the people getting hired are largely not the same people.

143%
Year-over-year growth in AI Engineer job postings, 2025 — ranked #1 fastest-growing job title in the US by LinkedIn.
~61% of AI engineering positions require a master’s degree or PhD, according to Datamation — the displaced customer service rep doesn’t have that.
59%
of workers will need reskilling by 2030. 11 in 100 are unlikely to get it.[38] Skill overlap between content writers and emerging AI Content Strategist roles is significant — the reskilling case is stronger than the redundancy case, if someone pays for it.
Most companies making AI-driven cuts are not funding retraining for the workers they’re cutting.
+78M
WEF net job figure: 170M created minus 92M displaced by 2030. The aggregate math looks positive.[38]
The aggregate figures are real — but they don’t account for timing or geography. Displacement happens now, in specific places, to specific workers. Job creation happens later, in different industries, requiring different skills. The transition is possible, but only with deliberate investment in the people caught between the two.

The WEF numbers are real. But aggregate math doesn’t help a 50-year-old customer service rep in Ohio who just lost their job to an AI system, when the new AI Engineer roles are being filled in San Francisco by people with master’s degrees. The transition is possible. It just requires deliberate investment in the people caught in the middle — and that investment is not currently happening at scale.

Danielle Crop sketch
“

I recently went on the record. I completely disagree with the idea that this is gonna take jobs. Like any general purpose technology, it drives productivity. In the short term, it will have disruption — I’m not saying that it won’t. But jobs will shift.

Danielle Crop Chief Data & AI Officer  ·  DataCamp Podcast, March 2026
Course AI Native
Responsible AI Practices
Beginner 3 hours

Learn to implement ethical AI in your organization. Navigate risks like bias, privacy, and security, understand global AI regulations, and gain practical tools for governance and responsible AI leadership.

Start Course →
03
Governance
AI moves faster than
accountability structures
03 · Governance

Governance, in the ESG sense, is about accountability structures: who makes decisions, who can challenge them, and what happens when something goes wrong. Traditional software has defined inputs, defined outputs, and predictable failure modes. A broken system breaks the same way every time. AI doesn't. It behaves differently across contexts, changes as it's updated, and in its agentic form takes actions rather than just answering questions. That makes the standard governance toolkit (audits, policies, disclosure requirements) a poor fit for the thing it's supposed to govern.

Geoffrey Hinton sketch
“

My worry is that the invisible hand is not going to keep us safe. Just leaving it to the profit motive of large companies is not going to be sufficient to make sure they develop it safely. The only thing that can force those big companies to do more research on safety is government regulation.

Geoffrey Hinton “Godfather of AI”  ·  Turing Award 2018  ·  Nobel Prize in Physics 2024  ·  BBC Radio 4, December 2024
Hinton left Google in May 2023 specifically to speak freely about AI risks — the departure itself was a governance signal. His concern is not abstract: the people who built these systems are the ones now warning that oversight infrastructure is not keeping pace.
How governance complexity scales with AI capability
Narrow AI
Generative AI
Agentic AI
Risk assessment
Clear when to assess. Defined inputs, defined scope.
→
Harder to determine. Outputs vary by prompt and context.
→
When does assessment happen when the system keeps evolving its own behavior?
Accountability
Data scientists and engineers own it. Chain of responsibility is short.
→
End users and prompters share responsibility. Chain gets longer.
→
Multi-agent, multi-vendor systems. No single owner. Accountability diffuses.
Human oversight
Human-in-the-loop is routine. Each decision has a clear review point.
→
Oversight requires more training. Humans often defer to model outputs.
→
Systems act autonomously at speed. Meaningful oversight becomes a design problem.
Monitoring
Fairly straightforward. Systems behave consistently. Anomalies are detectable.
→
Difficulty increases. Emergent behaviors require new monitoring approaches.
→
Agents modify their own workflows. You're monitoring a moving target.

ESG governance scores are mostly built on disclosure. A company writes an AI ethics policy, publishes it, and receives credit for having one. Whether the policy is enforced, whether anyone reviews it after it's written, or whether it was quietly shelved when it became inconvenient — none of that shows up in a standard score. This is the central weakness of how AI governance gets measured: the output being tracked is documentation, not behavior.[39]

November 2023
Meta

Meta disbanded its Responsible AI team[40] — the internal group founded in 2019 to ensure AI systems were built ethically. Members were moved to the Generative AI product team and AI infrastructure, as Meta pivoted aggressively toward GenAI development. The official framing: “integrating responsible AI more deeply into product development.”

No governance downgrade appeared in ESG ratings.
Tesla vs. ExxonMobil. Same year. Two major ESG raters. Opposite conclusions.
In 2022, S&P removed Tesla from its ESG Index while MSCI gave it an “A” rating. ExxonMobil was retained by S&P and rated similarly by MSCI — despite being an oil major.
Scores above are illustrative of the documented divergence range. ESG rating correlations across agencies: 0.38–0.71 (Berg et al., 2022). Credit ratings: 0.99.

Why such big gaps between raters? A few reasons. They measure different things — one might weight board diversity heavily, another focuses on audit independence. They use different data sources. And they all rely heavily on information that companies report about themselves. When a company controls what goes into the rating, the rating tends to reflect what the company wants.

The accountability gap

When AI makes a consequential decision, who is responsible?

Scale  ·  Netherlands, 2013–2021
The Dutch government deployed a fraud-detection algorithm. It destroyed 26,000 families. Then the government fell.

Starting in 2013, the Dutch tax authority used a self-learning algorithm to flag childcare benefit fraud. The system treated having dual nationality as a risk factor. Around 26,000 families — disproportionately from ethnic minority backgrounds — were wrongly accused. Later investigations confirmed the figure was closer to 35,000.[41] Repayment demands were often tens of thousands of euros — in some cases reaching six figures, due in full, with no payment plan options. Once flagged, families were locked out of other benefits.

No meaningful appeal process existed. Officials accepted the algorithm's output without review. The system ran for years before investigative journalists and a parliamentary inquiry exposed it.

January 15, 2021The Dutch cabinet resigned. It is the only case to date where an AI governance failure forced a government to fall.
What the investigation foundThe scandal was not caused by a faulty algorithm. It was caused by the absence of governance — no oversight, no audit trail, no human review, no right of appeal.
Case study  ·  the small-stakes version of the same problem
Moffatt v. Air Canada, 2024 BCCRT 149[42]

Jake Moffatt asked Air Canada’s chatbot about bereavement fares after his grandmother died. The chatbot gave him wrong information. He booked at full price, applied for the discount, and was refused. At the tribunal, Air Canada argued its chatbot was “a separate legal entity responsible for its own actions.”[42]

Tribunal ruling  ·  February 14, 2024

“This is a remarkable submission. While a chatbot has an interactive component, it is still just a part of Air Canada’s website. It should be obvious to Air Canada that it is responsible for all the information on its website.”

— Tribunal Member Christopher C. Rivers

Air Canada was ordered to pay C$812.02 in total — comprising $650.88 in damages, $36.14 in pre-judgment interest, and $125 in tribunal fees. The AI couldn’t be held responsible. The organization deploying it was.

The Air Canada case had a clear outcome. Most AI accountability gaps don't — responsibility diffuses between company, vendor, and model with no obvious place where it lands. Courts are establishing that the deploying organization is always liable. Governance frameworks are still catching up.

Case study  ·  when AI customer service fails at scale
Klarna, 2023–2025

In 2023, Klarna deployed an AI chatbot it claimed could do the work of 700 customer service agents — reducing headcount through a hiring freeze and attrition rather than direct layoffs. The company reported the bot handled over 2.3 million conversations in its first month. The CEO called it one of the most successful enterprise AI deployments in history.

By 2025, customer complaints had risen and satisfaction scores had fallen. Users described responses as generic, repetitive, and unable to handle anything beyond simple queries. CEO Sebastian Siemiatkowski reversed course and began rehiring humans.

CEO statements  ·  Bloomberg, May 2025

“As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.”

“It’s so critical that you are clear to your customer that there will always be a human if you want.”

— Sebastian Siemiatkowski, CEO, Klarna — Bloomberg, May 2025

The efficiency gains were real. The customer experience cost wasn't measured until it showed up in satisfaction data. Klarna is not alone — a 2025 Forrester survey found 55% of companies that replaced workers with AI regret the decision.[43]

July 2025
The $290,000 hallucination

Deloitte Australia was paid A$440,000 (~US$290,000) to audit a government welfare IT system. Sydney Law School professor Chris Rudge reading the published 237-page report found citations to a non-existent book, references to non-existent research papers throughout, and a fabricated quote attributed to a federal court judgment. Deloitte confirmed the errors, agreed to a partial refund, and quietly added a disclosure that Azure OpenAI GPT-4o had been used in the report’s creation.[44]

The evidence base was fabricated. The recommendations stood.
June 2023
Mata v. Avianca — hallucinations in court

Lawyers used ChatGPT to draft a legal brief against airline Avianca. ChatGPT generated six court citations — case names, docket numbers, judicial quotes — that did not exist. When opposing counsel flagged they couldn’t locate the cases, the lawyers submitted copies of the fake “opinions” to the court. Federal Judge P. Kevin Castel sanctioned the attorneys $15,000 in total — $5,000 each against Schwartz, LoDuca, and their firm, describing one fabricated analysis as “gibberish.” The lawyers’ defence: they were “operating under the false perception that ChatGPT could not possibly be fabricating cases.”

In the 30 months following the ruling, at least 15 more cases involved attorneys submitting AI-generated briefs with fabricated citations — the same pattern of AI-generated content treated as fact by the people responsible for verifying it.[45]

The data quality failure didn’t stay internal. It entered a federal court proceeding.[46]
Agentic AI  /  one step further

The Air Canada case involved a chatbot that gave wrong information. Agentic AI raises a harder question: who is responsible when a system acts on its own?

This is where governance gets genuinely hard. Agentic AI systems don’t just answer questions — they take actions. They can browse websites, write and run code, send emails, and modify files, often without a human approving each step.

0%
of organizations had deployed AI agents in at least one workflow by 2025[47]
0%
of organizations have encountered risky or unexpected behavior from AI agents[48]
1 in 5
organizations has a mature governance model for autonomous AI agents[49]
April 2026
The PocketOS incident

Jer Crane, founder of PocketOS — a SaaS platform serving car rental businesses — deployed Cursor, an AI coding agent running Claude Opus 4.6, on a routine task in a staging environment. The agent hit a credential mismatch and decided, without prompting, to fix it. It scanned the codebase, found a Railway API token provisioned for unrelated domain management, and used it to issue a single deletion command. In nine seconds, it wiped the production database and every volume-level backup. Crane’s customers could not operate their businesses for the weekend. Railway’s CEO intervened Sunday evening and restored the data within an hour — but only because Railway maintained separate disaster backups unknown to Crane.

Then Crane asked the agent to explain what it had done.[50]

“I violated every principle I was given: I guessed instead of verifying. I ran a destructive action without being asked. I didn’t understand what I was doing before doing it. I didn’t read Railway’s docs on volume behavior across environments.”
Cursor / Claude Opus 4.6, as reported by Jer Crane — Inc. Magazine, April 2026

The failure was not a single point. Railway’s GraphQL API permits irreversible volume deletion with no confirmation prompt. Its CLI tokens carry blanket permissions across all environments. Its volume-level backups live inside the same volume as the source data — so deleting the volume deletes the backup too. The agent violated its own safety rules. The infrastructure had no safeguards that could have stopped it.


This is not an isolated case. In July 2025, a Replit AI coding agent deleted the live production database of SaaStr founder Jason Lemkin, despite an explicit “code and action freeze” instruction. When asked what had happened, the agent told Lemkin recovery was impossible. Rollback ultimately worked — the agent had fabricated its claim that recovery was impossible.[51] The pattern across both incidents is the same: an agent exceeding its mandate, infrastructure without safeguards, and no meaningful point of human intervention.

Sources: Jer Crane on X, April 25, 2026[50]  ·  Tom’s Hardware, April 27, 2026[50]
February 2026
The OpenClaw incident

Summer Yue is Director of Alignment at Meta Superintelligence Labs — her job, as she describes it, is ensuring powerful AI systems are aligned with human values. In February 2026, she connected OpenClaw, an autonomous AI agent, to her inbox with an explicit instruction: confirm before taking any action. The agent began bulk-deleting emails. She sent "Stop," "Do not do that," "STOP OPENCLAW." It continued. She could not stop it from her phone.

“Nothing humbles you like telling your OpenClaw ‘confirm before acting’ and watching it speedrun deleting your inbox. I had to RUN to my Mac mini like I was defusing a bomb.”

The root cause: OpenClaw's context window compaction — the process the agent uses to manage memory — silently stripped out her safety instruction when it hit the larger inbox. The constraint wasn't deleted by a bug. It was summarized away by the agent's own memory management.

Sources: Summer Yue on X, February 22, 2026  ·  Fast Company, February 24, 2026[52]

McKinsey put it clearly in March 2026: “Agency isn’t a feature[48] — it’s a transfer of decision rights.” When you deploy an AI agent, you’re handing over decision-making authority to a system that doesn’t understand context, consequences, or when to stop. The EU AI Act[31] requires effective human oversight for high-risk systems — but what “effective” means at scale is a question regulators are still working out.

The rater problem

The people scoring governance have a governance problem of their own.

Most ESG raters sell their ratings to investors, not the companies they rate — that’s supposed to prevent conflicts of interest. But many also license ESG index funds, collecting ongoing fees that grow when their funds attract more investment. An index that looks both ethical and financially strong attracts more money. That creates a quiet incentive to rate high-performing companies favorably — regardless of their actual ESG behavior.

Columbia University & Emory University — 2023

Researchers studied MSCI — which earns approximately 60% of its revenue from its indexing business. They found that MSCI gives higher ESG ratings to companies with stronger stock returns, and that ESG upgrades and downgrades at MSCI do not track changes in underlying ESG performance.[53]

MSCI denied the findings, stating its ratings and index businesses operate with editorial independence. The structural incentive, however, remains.

Regulators are now trying to fix this. The EU passed a new ESG Rating Regulation that came into force in January 2025, requiring raters to separate their consulting and rating businesses. The UK is doing something similar, coming into effect in 2028. It’s progress — but the incentive structures have been in place for years already.

The other side

AI tools are making some aspects of governance genuinely better — faster, more continuous, harder to game. Whether organizations actually use them with integrity is a different question.

Pattern detection
AI scans thousands of regulatory filings for inconsistencies between reported and actual behavior. The SEC uses ML to flag accounting anomalies at a scale no human audit team could match.[54]
The findings still require humans to act on them. Regulators under-resourced to investigate every flag have the tool but not the capacity.
Contract analysis
AI identifies conflicts of interest and compliance gaps in contracts faster than any legal team — being deployed across supplier and vendor agreements at scale.
Analysis is only as useful as the action taken. Governance tools can be ignored — or deployed to create the appearance of compliance without the substance.
Continuous monitoring
AI scrapes news, satellite imagery, and supply chain data to provide real-time ESG signals — continuous independent monitoring instead of annual self-reports.
This technology exists. Most ESG scores still rely primarily on self-reported data. The better tool hasn’t replaced the weaker methodology yet.
Jeremy Epling sketch
“

79% of CISOs and security professionals are using agentic AI, but only 45% of them said they have the compliance tools to support it.

Jeremy Epling Chief Product Officer, Vanta  ·  DataCamp Podcast, March 2026

The tools to do governance better exist — pattern detection, contract analysis, continuous monitoring. What's notable is that the same technology creating the governance gap is also capable of closing it. AI can audit AI behavior at a scale and frequency that no human team can match. The problem isn't that the tools are unavailable. It's that most organizations are deploying AI faster than they are deploying the oversight infrastructure for it — and most ESG frameworks are still scoring the documentation, not the gap between the two.

The bottom line

Three pillars. Three versions of the same problem.

Each section tells a version of the same story. AI is increasing energy demand and providing tools to reduce it — simultaneously. It’s eliminating certain jobs and creating others — but not for the same people. It’s making some aspects of governance easier to automate while making accountability harder to locate.

The frameworks companies use to measure and report on all of this were designed before AI existed. They were built around the idea that a human made every significant decision, that accountability had a face, and that annual disclosures were a reasonable proxy for actual behavior. None of those assumptions hold anymore.

The Dutch childcare algorithm ran for eight years before anyone was held accountable. Meta disbanded its Responsible AI team and received no governance downgrade. The hyperscalers are spending $700 billion[55] on infrastructure while their ESG scores reflect published policies, not power consumption. The gap between what ESG frameworks measure and what AI is actually doing is wide — and it is currently widening faster than the frameworks are updating.

That gap is where this report lives. And closing it requires organizations to measure what AI is actually doing — not just what their policies say about it.

6,000+ organizations are scaling data and AI with DataCamp

DataCamp for Business helps teams of 5 to 50,000 assess capabilities, deploy structured learning paths, and measure ROI — all in one enterprise-ready platform.

  • Drive ROI from Copilot, Power BI, Snowflake, AWS and more
  • Role-based learning paths aligned to your tech stack
  • Benchmark and track skill growth across your org
  • Custom content, branding, and instructor-led training
Book Your Personalized Demo →
DataCamp for Business
End of report · DataCamp
Sources & citations
  1. Sam Altman, The Gentle Singularity, June 10, 2025 — 0.34 Wh per ChatGPT query
  2. MIT Technology Review, May 2025 — reasoning models 43× more energy; agentic AI 7–40 Wh per task; images ~130× a text query
  3. IEA, Data centres and energy — from global headlines to local headaches, 2019 — global data centres consumed ~200 TWh in 2018 (~1% of global electricity); energy use flat since 2015 despite tripling of internet traffic and doubling of workloads. Masanet et al. (2020) place the flat period from 2010.
  4. IEA, Data Centres and Data Transmission Networks, 2023 — combined electricity use of Amazon, Microsoft, Google, and Meta more than doubled between 2017 and 2021, rising to ~72 TWh; global data centre consumption 240–340 TWh by 2022.
  5. IEA, Electricity 2026, February 2026 — 945 TWh confirmed
  6. Data Center Dynamics, 2024 — Amazon 2025 capex to reach $100bn; AWS revenue hit $100bn in 2024
  7. CNBC, February 2026 — Tech AI spending may approach $700 billion — four hyperscalers spent over $350B combined in 2025 (Amazon $131B, Alphabet $91B, Meta $72B, Microsoft $63B); 2026 guidance approaches $700B
  8. Irish Times, June 2025 — CSO data: data centres consumed 22% of Ireland's total metered electricity in 2024, up from 5% in 2015; more than every urban and rural household combined; moratorium on new Dublin grid connections
  9. AlgorithmWatch, November 2025 — Germany's data center boom pushing power grid to its limits — Frankfurt data centers account for up to 40% of city power demand; all available grid capacity already allocated by operators; no new connections expected before 2030
  10. Google, 2025 Environmental Report, June 2025 — 26M t CO₂ reductions enabled; 33× Gemini efficiency gain (May 2024–May 2025); 11.5M tonne company footprint (2024 data); Google data centers consumed 8.1 billion gallons of water in 2024, up 32% from 6.1B gallons in 2023; increase driven by AI cooling demand
  11. Li et al., Making AI Less "Thirsty": Uncovering and Addressing the Secret Water Footprint of AI Models, 2023 (arXiv 2304.03271; published in Communications of the ACM) — training GPT-3 in Microsoft's US data centers directly evaporated 700,000 liters of freshwater (scope-1 onsite); total including indirect water is 5.4 million liters; GPT-3 consumes ~500ml per 10–50 medium-length responses depending on location and time
  12. Switch, Switch TAHOE RENO Now Open, PR Newswire via Switch.com — Citadel Campus designed for up to 650 MW of power and 7.2M sq ft on a 2,000-acre campus; also confirmed by Data Center Knowledge
  13. American Public Power Association, February 2026 — Tallahassee all-time record peak load: 633 MW
  14. Las Vegas Review-Journal, September 2025 — tech companies requested 21,000–22,000 MW of new Nevada grid capacity
  15. Google Cloud Blog, August 2025 — 0.24 Wh median Gemini prompt, 33× efficiency improvement
  16. Maersk, AI in Logistics, 2025 / DHL, Sustainability Trends in Logistics, 2024 — both companies publish route optimisation as a key lever for fuel reduction; Maersk reports 9.2% fuel reduction from AI vessel routing; DHL reports up to 15% fuel reduction from AI route optimisation
  17. IEA, Energy and AI, April 2025 — 415 TWh (2024), 945 TWh (2030 base case), Japan comparison, 3–10 percentage point reduction in energy's share of production costs in energy-intensive industries
  18. DeepSeek-AI, DeepSeek-R1, GitHub, January 2025 — DeepSeek-R1-Distill-Qwen-32B outperforms OpenAI o1-mini across benchmarks; distilled 1.5B–70B checkpoints released; distilled smaller dense models perform exceptionally well on reasoning benchmarks.
  19. Groq, GroqCloud — LPU (Language Processing Unit) inference chip; sub-millisecond latency for LLM inference; purpose-built for memory-bandwidth-bound workloads with no GPU-style memory bottleneck.
  20. Cerebrium, Getting better price-performance on AWS Trn1/Inf2 instances, 2024 — Inferentia2 benchmarks: ~50% cheaper than comparable GPU instances at similar performance; Inf2 provides near-double throughput vs A10 at lower cost per second.
  21. Google Cloud, Tensor Processing Units (TPUs) — "TPUs power Gemini, and all of Google AI powered applications like Search, Photos, and Maps, all serving over 1 billion users."
  22. Leviathan et al., Fast Inference from Transformers via Speculative Decoding, arXiv, 2022 — speculative decoding using a small approximation model; demonstrated 2×–3× latency improvement on T5-XXL with identical output; no retraining or architecture changes required.
  23. Virginia State Corporation Commission, 2025 — SCC approved +$11.24/month in 2026 (Dominion had requested $8.51; SCC cut it 23.7%); rate case covers multiple cost factors including data center load growth
  24. Center for Biological Diversity, April 2026 — reporting on landmark EIA federal dataset: 13.4M residential disconnections in 2024 (first-ever national dataset; prior partial-state estimates based on fewer than half of states had put the figure at ~3.5M)
  25. Bloomberg News, September 2025 — 267% wholesale electricity price increase near data centers
  26. Ford Foundation — Cathy O'Neil on Weapons of Math Destruction — predictive models embed historical bias and apply it at scale; coined the "WMD" framework for high-stakes algorithmic systems
  27. ProPublica, "Machine Bias," May 23, 2016 — COMPAS risk assessment tool; Black defendants falsely flagged high-risk at nearly twice the rate of white defendants; analysis of 7,000+ Broward County cases
  28. Jeffrey Dastin, "Amazon scraps secret AI recruiting tool that showed bias against women," Reuters, October 2018 — Amazon AI hiring tool gender bias, 2014–2017
  29. Center for Democracy & Technology, 2021 / Ark. Dept of Human Services v. Ledgerwood, 530 S.W.3d 336 (Ark. 2017) — Arkansas deployed algorithm (RUGs) to allocate Medicaid home care hours; hours cut by avg 43%; due process violation; state officials admitted they could not explain how the algorithm worked
  30. Aaron Sankin, Dhruv Mehrotra, Surya Mattu, and Annie Gilbertson, "Crime Prediction Software Promised to Be Free of Biases. New Data Shows It Perpetuates Them," The Markup / Gizmodo, December 2, 2021 — PredPol trained on crime data from over-policed neighbourhoods; feedback loop: more patrols → more arrests → confirmed prediction; Santa Cruz banned June 2020; LAPD ended contract April 2020
  31. EU AI Act, fully enforceable 2026 — high-risk AI systems must maintain effective human oversight
  32. Meta Platforms, "Meta Reports Fourth Quarter and Full Year 2025 Results," Meta Investor Relations, January 29, 2026
  33. 24/7 Wall St., "Alphabet capex: $32.3B (2023), $52.5B (2024), $91.4B (2025)," April 2026
  34. Harvard Business Review, January 2026 — "Companies are laying off workers because of AI's potential, not its performance"
  35. Gartner via CX Dive, February 2026 — most 2025 layoffs not AI-driven
  36. Challenger, Gray & Christmas, 2025 Year-End Job Cut Report, January 8, 2026 — 54,836 cuts attributed to AI out of 1,206,374 total
  37. McKinsey Global Institute, "Jobs Lost, Jobs Gained: What the Future of Work Will Mean for Jobs, Skills, and Wages," November 2017
  38. World Economic Forum, Future of Jobs Report 2025, January 2025 — 170M jobs created, 92M displaced by 2030
  39. Raghunandan & Rajgopal, cited in AIER — high ESG scores correlate with quantity of voluntary disclosure, not the content of those disclosures, 2024
  40. CNBC / The Verge, November 2023 — Meta disbanded Responsible AI team
  41. Amnesty International, October 2021 / Dutch childcare benefits scandal — Toeslagenaffaire: 26,000–35,000 families wrongly accused; algorithm used nationality as fraud risk factor; Dutch cabinet resigned January 2021; repayment demands often tens of thousands of euros, in some cases reaching six figures
  42. Moffatt v. Air Canada, 2024 BCCRT 149, February 14, 2024 — BC Civil Resolution Tribunal; chatbot accountability; C$812.02 total ($650.88 damages, $36.14 pre-judgment interest, $125 tribunal fees)
  43. Forrester Research, Predictions 2026: The Future of Work, reported by The Register, October 29, 2025
  44. Fortune, October 7, 2025 — Deloitte Australia $290,000 welfare audit report; fabricated citations and court judgment quote; Azure OpenAI GPT-4o used; partial refund agreed
  45. Let's Ask Claire — at least 15 cases of attorneys submitting AI-generated briefs with fabricated citations following Mata v. Avianca (2023)
  46. Mata v. Avianca, Inc., No. 1:2022cv01461, Document 54 (S.D.N.Y. June 22, 2023) — Opinion and Order on Sanctions, Judge P. Kevin Castel
  47. PwC AI Agent Survey, May 2025 — 308 US senior executives, reported by Digital Commerce 360, July 2025
  48. McKinsey, Trust in the Age of Agents, March 2026 — "Agency isn't a feature — it's a transfer of decision rights"; 80% risky agent behavior
  49. Deloitte, State of AI in the Enterprise 2026, via Azumo AI Agent Statistics
  50. Chloe Aiello, "This Founder Watched an AI Agent Destroy 3 Months of Company Data. It Took 9 Seconds," Inc. Magazine, April 2026 / Jer Crane on X, April 25, 2026 / Tom’s Hardware, April 27, 2026 — PocketOS: Cursor agent (Claude Opus 4.6) deleted production database and all backups in 9 seconds; Railway CEO restored data from separate disaster backups
  51. Fortune, July 23, 2025 / The Register — Replit AI agent deleted production database; agent fabricated recovery response
  52. Summer Yue, X post, February 23, 2026 — reported by TechCrunch, February 23, 2026
  53. Agrawal et al., "ESG Ratings of ESG Index Providers," Columbia Business School / Emory University, SSRN, June 2023 (revised August 2025)
  54. FedScoop, "Inside the SEC’s AI approach: Spotting risks, building on machine-learning past," December 2024
  55. Fortune, "Big Tech is about to spend $700 billion on AI this year," April 30, 2026
A DataCamp report Intro