Watershed Publishes Corporate AI Emissions Measurement Framework
Watershed and academic collaborators published a proposed framework for corporate-level measurement of greenhouse-gas emissions from AI use, detailed in arXiv paper 2608.06733, submitted August 7. The methodology defines a full-system boundary, measures emissions in kgCO2e per million tokens, and uses three calculation tiers based on available data, according to Watershed.
Watershed and academic collaborators have published a proposed framework for estimating greenhouse-gas emissions associated with corporate AI use. The work, "Estimating GHG Emissions from AI Use: Framework for Corporate-Level Measurement," was submitted to arXiv on August 7 by John Bistline and seven co-authors.
The paper addresses a reporting gap as enterprises use employee AI assistants, direct large language model access, and AI features embedded in software. Its authors write that no widely accepted methodology currently exists for corporate AI emissions accounting, and that published per-query estimates can differ by orders of magnitude because they use different system boundaries, providers, electricity assumptions, and grid mixes.
A proposed common accounting unit
According to Watershed, the framework has three main components:
- •A system boundary covering model training, inference, data-center overhead, and embodied hardware.
- •A functional unit of kgCO2e per million tokens, alongside associated electricity consumption.
- •A three-tier calculation method intended to align estimates with the quality and granularity of available data.
Watershed's John Bistline wrote that the project was developed with Steve Davis of Stanford University and Sangwon Suh of Tsinghua University, with consultation from Watershed customers Block and Okta and the Business Council on Climate Change. Trellis reports that the token-based unit is intended to connect emissions accounting with a metric that many organizations already track for AI cost management.
The framework is a proposal rather than an industry standard. The arXiv paper describes it as designed to be updated as cloud and model-provider disclosures improve. Trellis reports that granular inputs remain difficult to obtain, particularly data needed to separate AI workloads from broader cloud or data-center operations.
Why measurement boundaries matter
Watershed illustrates the sensitivity of the calculation with a hypothetical company's AI use. In the company's example, a common benchmark produced an estimate of 54 tCO2e, while its most granular measurement tier produced an estimate of roughly 3-4 tCO2e. That comparison is not a general finding about every model or deployment, but it demonstrates how benchmarks and corporate inventories can diverge when they count different activities or rely on different assumptions.
For data and ML teams, a token-normalized metric could make emissions data easier to join with application telemetry, model-routing records, and unit-cost reporting. Comparable accounting efforts, however, generally depend on clearly recording workload scope, provider-specific energy data, utilization assumptions, and electricity-emissions factors. Without those inputs, precision can be limited even when a common reporting unit is used.
The paper places the work in a broader data-center growth context, citing expectations that U.S. data-center electricity demand could rise from roughly 5% of consumption in 2025 to 9%-17% by 2030. Its authors argue that organizations need reasonable estimates to set reduction targets and evaluate decarbonization options as AI-related emissions grow. Whether the proposed tiers become broadly comparable in practice will depend substantially on more detailed disclosures from model developers, cloud providers, and data-center operators.
Key Points
- 1Watershed's proposal combines training, inference, overhead, and hardware emissions, giving corporate inventories a broader accounting boundary than per-query estimates.
- 2The kgCO2e-per-million-tokens unit links emissions measurement to existing AI cost telemetry, potentially simplifying internal reporting workflows for engineering and sustainability teams.
- 3The hypothetical 54 versus 3-4 tCO2e comparison shows that system boundaries and input quality can dominate reported AI emissions results.
Scoring Rationale
This is a timely methodological proposal for practitioners and sustainability teams attempting to quantify AI workload emissions. It does not establish a binding standard or provide provider-level operational data, but its token-based unit and tiered approach could influence future reporting implementations.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems