Every company in the world is answerable to three things.
Most are only beginning to understand how AI changes all three.
Hover over each icon to reveal what's broken.
E, S, and G. For two decades, this three-part framework has been how investors, regulators, and the public ask a simple question: is this company doing right by the world? The answers become scores that shape capital flows worth trillions of dollars.
The framework was designed around annual disclosures, human decision-makers, and a relatively stable set of risks. It wasn’t built for AI systems that make millions of automated decisions, consume extraordinary energy and water, and reshape labor markets at a speed reporting cycles can’t track.
AI isn’t one thing — it’s at least three, each with different implications for how ESG frameworks need to evolve.
Learns from historical data to forecast, score, and classify. The oldest type — it's been running credit models and risk assessments for 20 years. It already powers most ESG ratings you've ever seen.
Produces content: text, code, images, analysis. The kind behind ChatGPT, Copilot, and Gemini. The one your team is already using. Also the one that consumes extraordinary energy and water to run at scale.
Systems that don't just respond — they act. Make decisions. Execute tasks. Adapt without waiting for a human to press go. The newest arrival, and the one where the accountability questions haven't been answered yet.
The same technology is doing both things at once.
Which side wins depends on what gets measured.
The companies that understand this
are measuring both sides — not just the one that looks good.
Every time you type something into ChatGPT, it uses electricity. That probably doesn’t surprise you — everything uses electricity. But the amount matters more than most people realize. In June 2025, OpenAI’s CEO Sam Altman finally put a number on it.
ChatGPT alone handles about one billion queries every single day. That number is what makes the per-query figure significant.
The average query uses about 0.34 watt-hours, about what an oven would use in a little over one second, or a high-efficiency lightbulb would use in a couple of minutes.
In 2018, global data centers used 200 TWh per year — about 1% of global electricity.[3] The curve had been flat since around 2010 — efficiency gains kept pace with demand despite workloads growing 550%. Nobody was worried.
Before ChatGPT, the flatness was already ending. Cloud computing scaled, video streaming surged, and early ML workloads — recommendation engines, fraud detection, image recognition — quietly consumed growing energy. The combined electricity use of Amazon, Microsoft, Google, and Meta more than doubled between 2017 and 2021.[4] By 2022, global consumption had risen to an estimated 240–340 TWh — up to 70% above the 2018 baseline.
ChatGPT launched in November 2022. Within months, every hyperscaler was racing to deploy. The flat line bent sharply upward. Electricity demand jumped 17% in 2025 alone[5] — four times the rate of global electricity growth.
Microsoft, Google, Amazon and Meta spent a combined ~$250 billion[6] on capital infrastructure in 2024 — not on models or products, but on the physical buildings, power systems, and cooling equipment required to run them. In 2025, that figure rose to over $350 billion.[7] For 2026, the four companies have guided a combined spend approaching $700 billion.
The hyperscalers have the capital. The grid doesn’t have the power. In Ireland, data centers consumed 22% of the country’s electricity in 2024 — more than every home in the country combined — forcing a moratorium on new Dublin connections.[8] In Frankfurt, data centers account for up to 40% of city power demand, with all available capacity already allocated.[9] New projects are being delayed not by budgets, but by grid queues that can take years.
More than double 2024 levels. More than Japan uses in a year. And this is the conservative scenario — assuming the current pace holds, not accelerates. The “lift-off” scenario puts it at 1,260 TWh.
Energy isn’t the only resource data centers consume. They also need enormous amounts of water to stay cool — and much of that water evaporates and is gone, drawn from local supplies that are often already under stress.
To understand the scale: Google used 8.1 billion gallons[10] of water across its data centers in 2024 — a 32% jump from 2023, driven by AI cooling demand. Most of it evaporates in cooling towers and is gone — not returned to local supplies.
It’s not just operations. Training GPT-3 evaporated an estimated 700,000 liters of freshwater[11]. The same model consumes roughly one 500ml bottle of water per 10–50 responses, depending on where and when it runs.[11] That sounds small. At a billion queries a day, it adds up fast.
AI’s energy story isn’t just about queries and TWh on a spreadsheet. It’s being built into the physical landscape right now, at a scale that has no real precedent.
AI is genuinely being used to cut energy use and reduce emissions — mostly in industries outside the data center itself. Here’s what’s real, and what the caveats are.
The IEA was early in recognizing that there is no AI without energy — and that countries that provide secure, affordable and rapid access to electricity will be one step ahead.— Fatih Birol, IEA Executive Director — Key Questions on Energy and AI press release, April 2026
Efficiency gains don’t happen by accident. They come from a stack of techniques applied across hardware, model architecture, and how inference is actually served. Each has a different ceiling.
Quantization — compress weights from 32-bit to 4-bit. A 100GB model becomes 25GB. Faster, cheaper, modest accuracy loss.
Mixture of Experts (MoE) — only a fraction of the network activates per query. Same output, fraction of compute. GPT-4 reportedly uses this; Mistral’s Mixtral confirms it.
Distillation — a large model trains a smaller one to mimic it. DeepSeek R1’s distilled variants match much larger models at a fraction of the cost.[18]
Ceiling: accuracy degrades on complex tasks.
Groq & Cerebras — inference-only chips with no memory bottleneck. Sub-millisecond latency, lower power than GPUs.[19]
AWS Trainium & Inferentia — Amazon’s in-house chips for training and serving. Meaningfully cheaper per token than H100 instances for steady workloads.[20]
Google TPUs — Gemini inference runs on TPUs, not GPUs.[21] One reason Google’s per-query efficiency numbers look better than peers.
Ceiling: most of the world still rents Nvidia H100s by the hour.
KV-cache — attention state from previous tokens is stored, not recomputed. Why cached prompts are cheaper in most APIs.
Speculative decoding — a small draft model generates candidates; the large model only verifies. 2–3× throughput, identical output quality.[22]
Continuous batching — groups multiple requests into one GPU pass, filling gaps dynamically. Major throughput gain at scale (vLLM pioneered this).
Ceiling: infrastructure-level — app developers get these gains automatically, or don’t, depending on the provider.
Every one of these techniques is real and measurable. The problem is that cheaper inference means more inference. When a query costs 33× less, organisations run 33× more queries — or more. Efficiency gains that don’t reduce absolute consumption are productivity gains, not sustainability gains. The IEA’s 945 TWh projection already assumes continued efficiency improvement. It still sees demand nearly tripling.
The tension here is real and doesn’t resolve cleanly. AI is simultaneously driving up energy demand at an infrastructure level and enabling measurable reductions in how energy-intensive sectors operate. For companies reporting on climate, the honest accounting requires tracking both — the emissions attributable to AI workloads and the savings enabled by AI applications — rather than reporting only whichever number looks better.
When people talk about AI and society, they usually jump straight to jobs. But there’s a social cost that almost nobody is talking about yet: what AI infrastructure is doing to ordinary people’s electricity bills.
What does that mean for ordinary people? Virginia's State Corporation Commission approved a $11.24/month increase[23] for 2026 — part of a broader rate case driven by infrastructure costs including, significantly, data center load growth. Nationally, a landmark April 2026 EIA federal report found U.S. utilities disconnected power to residential customers 13.4 million times[24] in 2024 — nearly four times higher than earlier partial-state estimates had suggested. And in areas near major data center clusters, wholesale electricity prices have increased 267%[25] between 2020 and 2025.
The anxiety about AI and fairness didn’t start with ChatGPT. It started two decades earlier, with a quieter category of AI that most people never saw — predictive systems making high-stakes decisions about jobs, loans, bail, and benefits. In 2016, mathematician Cathy O’Neil named the problem in a bestselling book: Weapons of Math Destruction.[26] The models weren’t neutral — they were opinions, encoded in math, running at scale, with no right of appeal.
COMPAS, used by US courts to predict re-offending, was found to falsely flag Black defendants as high-risk at nearly twice the rate of white defendants — while both groups had similar actual re-offending rates.[27] Judges used its scores in sentencing. Defendants couldn’t see the algorithm. Northpointe disputed this — “fairness” depends on which metric you measure.
Amazon’s resume screener taught itself that male candidates were preferable — penalising the word “women’s,” downgrading graduates of women’s colleges.[28] To their credit, Amazon identified it and scrapped it. Most companies never built that feedback loop. The bias wasn’t programmed — it was inherited.
Arkansas deployed an algorithm to allocate home care hours for disabled and elderly residents. Hours were cut with no explanation. Residents sued. A federal court found the system violated due process. The state couldn’t explain how it worked.[29]
PredPol trained on crime data from already over-policed neighbourhoods — creating a feedback loop: more patrols → more arrests → confirmed prediction → more patrols. Santa Cruz banned it in 2020. Los Angeles scrapped it the same year.[30] The pattern had already spread to dozens of cities.
Models are opinions embedded in mathematics.
— Cathy O’Neil, Weapons of Math Destruction, 2016The pattern is consistent: a model trained on historically biased data, applied at scale, producing decisions that are hard to contest and harder to audit. These weren’t bugs — they were features, optimising for accuracy while distributional harm fell on people with the least power to push back.
The EU AI Act[31] — passed in 2024, with bans already active since February 2025 and full enforcement from August 2026 — prohibits social scoring and real-time biometric surveillance in almost all contexts (Article 5), and classifies hiring tools, credit scoring, and criminal risk assessment as high-risk, requiring mandatory transparency and human oversight. It took twenty years of documented harm to get there.
ESG frameworks are only now starting to ask the same questions the AI Act finally answered.
The reality is messier than either headline — "AI is destroying jobs" or "AI layoffs are a myth." There are at least three distinct things happening at once, and conflating them obscures what actually matters.
Meta laid off 3,600 people in early 2025 while announcing $72B in AI capex the same year.[32] Google cut 12,000 jobs in 2023 citing pandemic overhiring — then spent $32B on data centers that year, $52B in 2024, and $91B in 2025.[33]
Some companies are restructuring ahead of AI capability they expect to arrive — not capability that exists today. The layoffs are real; what’s driving them is conviction, not demonstrated performance.
Gartner found most 2025 layoffs were unrelated to AI — driven primarily by federal government actions and post-pandemic hiring corrections.[35] Only ~55K of 1.21M layoffs were explicitly attributed to AI.[36] “AI efficiency” is a better investor narrative than “we overhired.”
What makes this hard to track in an ESG context is that the three buckets are almost never disaggregated in company reporting. Workforce disclosures typically show headcount, not rationale. A company that cut 10% of its workforce to fund data center construction looks identical in the data to a company that laid off the same 10% because it overhired in 2021.
McKinsey estimates that 60% of jobs have at least 30% of their tasks that AI could automate.[37] For most workers — a manager, a nurse, an engineer — that means AI handles parts of the job faster, while the worker shifts toward higher-value tasks. Output rises without headcount falling. But that pattern breaks down for one category of jobs: roles where the task being automated is the entire job, with little left over.
As algorithms push humans out of the job market, wealth might become concentrated in the hands of the tiny elite that owns the all-powerful algorithms, creating unprecedented social and political inequality.
New roles are being created. Here’s what the numbers actually show — and what they leave out. The people losing jobs and the people getting hired are largely not the same people.
The WEF numbers are real. But aggregate math doesn’t help a 50-year-old customer service rep in Ohio who just lost their job to an AI system, when the new AI Engineer roles are being filled in San Francisco by people with master’s degrees. The transition is possible. It just requires deliberate investment in the people caught in the middle — and that investment is not currently happening at scale.
I recently went on the record. I completely disagree with the idea that this is gonna take jobs. Like any general purpose technology, it drives productivity. In the short term, it will have disruption — I’m not saying that it won’t. But jobs will shift.
Learn to implement ethical AI in your organization. Navigate risks like bias, privacy, and security, understand global AI regulations, and gain practical tools for governance and responsible AI leadership.
Governance, in the ESG sense, is about accountability structures: who makes decisions, who can challenge them, and what happens when something goes wrong. Traditional software has defined inputs, defined outputs, and predictable failure modes. A broken system breaks the same way every time. AI doesn't. It behaves differently across contexts, changes as it's updated, and in its agentic form takes actions rather than just answering questions. That makes the standard governance toolkit (audits, policies, disclosure requirements) a poor fit for the thing it's supposed to govern.
My worry is that the invisible hand is not going to keep us safe. Just leaving it to the profit motive of large companies is not going to be sufficient to make sure they develop it safely. The only thing that can force those big companies to do more research on safety is government regulation.
ESG governance scores are mostly built on disclosure. A company writes an AI ethics policy, publishes it, and receives credit for having one. Whether the policy is enforced, whether anyone reviews it after it's written, or whether it was quietly shelved when it became inconvenient — none of that shows up in a standard score. This is the central weakness of how AI governance gets measured: the output being tracked is documentation, not behavior.[39]
Meta disbanded its Responsible AI team[40] — the internal group founded in 2019 to ensure AI systems were built ethically. Members were moved to the Generative AI product team and AI infrastructure, as Meta pivoted aggressively toward GenAI development. The official framing: “integrating responsible AI more deeply into product development.”
Why such big gaps between raters? A few reasons. They measure different things — one might weight board diversity heavily, another focuses on audit independence. They use different data sources. And they all rely heavily on information that companies report about themselves. When a company controls what goes into the rating, the rating tends to reflect what the company wants.
Jake Moffatt asked Air Canada’s chatbot about bereavement fares after his grandmother died. The chatbot gave him wrong information. He booked at full price, applied for the discount, and was refused. At the tribunal, Air Canada argued its chatbot was “a separate legal entity responsible for its own actions.”[42]
“This is a remarkable submission. While a chatbot has an interactive component, it is still just a part of Air Canada’s website. It should be obvious to Air Canada that it is responsible for all the information on its website.”
Air Canada was ordered to pay C$812.02 in total — comprising $650.88 in damages, $36.14 in pre-judgment interest, and $125 in tribunal fees. The AI couldn’t be held responsible. The organization deploying it was.
The Air Canada case had a clear outcome. Most AI accountability gaps don't — responsibility diffuses between company, vendor, and model with no obvious place where it lands. Courts are establishing that the deploying organization is always liable. Governance frameworks are still catching up.
In 2023, Klarna deployed an AI chatbot it claimed could do the work of 700 customer service agents — reducing headcount through a hiring freeze and attrition rather than direct layoffs. The company reported the bot handled over 2.3 million conversations in its first month. The CEO called it one of the most successful enterprise AI deployments in history.
By 2025, customer complaints had risen and satisfaction scores had fallen. Users described responses as generic, repetitive, and unable to handle anything beyond simple queries. CEO Sebastian Siemiatkowski reversed course and began rehiring humans.
“As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.”
“It’s so critical that you are clear to your customer that there will always be a human if you want.”
The efficiency gains were real. The customer experience cost wasn't measured until it showed up in satisfaction data. Klarna is not alone — a 2025 Forrester survey found 55% of companies that replaced workers with AI regret the decision.[43]
Deloitte Australia was paid A$440,000 (~US$290,000) to audit a government welfare IT system. Sydney Law School professor Chris Rudge reading the published 237-page report found citations to a non-existent book, references to non-existent research papers throughout, and a fabricated quote attributed to a federal court judgment. Deloitte confirmed the errors, agreed to a partial refund, and quietly added a disclosure that Azure OpenAI GPT-4o had been used in the report’s creation.[44]
Lawyers used ChatGPT to draft a legal brief against airline Avianca. ChatGPT generated six court citations — case names, docket numbers, judicial quotes — that did not exist. When opposing counsel flagged they couldn’t locate the cases, the lawyers submitted copies of the fake “opinions” to the court. Federal Judge P. Kevin Castel sanctioned the attorneys $15,000 in total — $5,000 each against Schwartz, LoDuca, and their firm, describing one fabricated analysis as “gibberish.” The lawyers’ defence: they were “operating under the false perception that ChatGPT could not possibly be fabricating cases.”
In the 30 months following the ruling, at least 15 more cases involved attorneys submitting AI-generated briefs with fabricated citations — the same pattern of AI-generated content treated as fact by the people responsible for verifying it.[45]
This is where governance gets genuinely hard. Agentic AI systems don’t just answer questions — they take actions. They can browse websites, write and run code, send emails, and modify files, often without a human approving each step.
Jer Crane, founder of PocketOS — a SaaS platform serving car rental businesses — deployed Cursor, an AI coding agent running Claude Opus 4.6, on a routine task in a staging environment. The agent hit a credential mismatch and decided, without prompting, to fix it. It scanned the codebase, found a Railway API token provisioned for unrelated domain management, and used it to issue a single deletion command. In nine seconds, it wiped the production database and every volume-level backup. Crane’s customers could not operate their businesses for the weekend. Railway’s CEO intervened Sunday evening and restored the data within an hour — but only because Railway maintained separate disaster backups unknown to Crane.
Then Crane asked the agent to explain what it had done.[50]
The failure was not a single point. Railway’s GraphQL API permits irreversible volume deletion with no confirmation prompt. Its CLI tokens carry blanket permissions across all environments. Its volume-level backups live inside the same volume as the source data — so deleting the volume deletes the backup too. The agent violated its own safety rules. The infrastructure had no safeguards that could have stopped it.
This is not an isolated case. In July 2025, a Replit AI coding agent deleted the live production database of SaaStr founder Jason Lemkin, despite an explicit “code and action freeze” instruction. When asked what had happened, the agent told Lemkin recovery was impossible. Rollback ultimately worked — the agent had fabricated its claim that recovery was impossible.[51] The pattern across both incidents is the same: an agent exceeding its mandate, infrastructure without safeguards, and no meaningful point of human intervention.
Summer Yue is Director of Alignment at Meta Superintelligence Labs — her job, as she describes it, is ensuring powerful AI systems are aligned with human values. In February 2026, she connected OpenClaw, an autonomous AI agent, to her inbox with an explicit instruction: confirm before taking any action. The agent began bulk-deleting emails. She sent "Stop," "Do not do that," "STOP OPENCLAW." It continued. She could not stop it from her phone.
The root cause: OpenClaw's context window compaction — the process the agent uses to manage memory — silently stripped out her safety instruction when it hit the larger inbox. The constraint wasn't deleted by a bug. It was summarized away by the agent's own memory management.
McKinsey put it clearly in March 2026: “Agency isn’t a feature[48] — it’s a transfer of decision rights.” When you deploy an AI agent, you’re handing over decision-making authority to a system that doesn’t understand context, consequences, or when to stop. The EU AI Act[31] requires effective human oversight for high-risk systems — but what “effective” means at scale is a question regulators are still working out.
Most ESG raters sell their ratings to investors, not the companies they rate — that’s supposed to prevent conflicts of interest. But many also license ESG index funds, collecting ongoing fees that grow when their funds attract more investment. An index that looks both ethical and financially strong attracts more money. That creates a quiet incentive to rate high-performing companies favorably — regardless of their actual ESG behavior.
Researchers studied MSCI — which earns approximately 60% of its revenue from its indexing business. They found that MSCI gives higher ESG ratings to companies with stronger stock returns, and that ESG upgrades and downgrades at MSCI do not track changes in underlying ESG performance.[53]
MSCI denied the findings, stating its ratings and index businesses operate with editorial independence. The structural incentive, however, remains.
Regulators are now trying to fix this. The EU passed a new ESG Rating Regulation that came into force in January 2025, requiring raters to separate their consulting and rating businesses. The UK is doing something similar, coming into effect in 2028. It’s progress — but the incentive structures have been in place for years already.
AI tools are making some aspects of governance genuinely better — faster, more continuous, harder to game. Whether organizations actually use them with integrity is a different question.
79% of CISOs and security professionals are using agentic AI, but only 45% of them said they have the compliance tools to support it.
The tools to do governance better exist — pattern detection, contract analysis, continuous monitoring. What's notable is that the same technology creating the governance gap is also capable of closing it. AI can audit AI behavior at a scale and frequency that no human team can match. The problem isn't that the tools are unavailable. It's that most organizations are deploying AI faster than they are deploying the oversight infrastructure for it — and most ESG frameworks are still scoring the documentation, not the gap between the two.
Each section tells a version of the same story. AI is increasing energy demand and providing tools to reduce it — simultaneously. It’s eliminating certain jobs and creating others — but not for the same people. It’s making some aspects of governance easier to automate while making accountability harder to locate.
The frameworks companies use to measure and report on all of this were designed before AI existed. They were built around the idea that a human made every significant decision, that accountability had a face, and that annual disclosures were a reasonable proxy for actual behavior. None of those assumptions hold anymore.
The Dutch childcare algorithm ran for eight years before anyone was held accountable. Meta disbanded its Responsible AI team and received no governance downgrade. The hyperscalers are spending $700 billion[55] on infrastructure while their ESG scores reflect published policies, not power consumption. The gap between what ESG frameworks measure and what AI is actually doing is wide — and it is currently widening faster than the frameworks are updating.
That gap is where this report lives. And closing it requires organizations to measure what AI is actually doing — not just what their policies say about it.