Why can't we just measure training and call it done?
Most public attention has focused on the energy used to train large AI models, but that is only part of the picture. Training is a one-time cost, while inference—the act of running a model to answer queries—happens millions or billions of times once a model is deployed. A 2023 study on generative AI inference found that for a service like ChatGPT, the cumulative energy and carbon impact of inference can quickly surpass training, especially as user numbers grow [7]. This means that if you only measure training, you could dramatically underestimate the true footprint of an AI system.
Several reviews confirm this blind spot. A 2026 scoping review notes that while training has been studied extensively, the 'cumulative environmental footprint generated during large-scale service operations, particularly in the inference phase, has received comparatively less attention' [3]. Another 2025 review introduces a six-dimensional framework that explicitly includes both training and deployment (inference) as separate but linked stages [2]. The bottom line: training alone is an incomplete metric.
What do the studies agree on?
There is broad consensus across the papers that a standardized, transparent measurement framework covering the full AI lifecycle is essential—and currently missing. A 2026 paper on language models found that 'existing benchmarks and publications rarely report energy consumption, creating a significant information gap' [1]. Similarly, a 2023 guide to carbon footprint tools notes that several online calculators exist, but they produce wildly different estimates: one comparison of four calculators found that for the same AI query, estimates varied by a factor of more than 50 [6]. This makes it nearly impossible for users to compare models or make informed choices.
Another point of agreement is that hardware choices and model architecture matter enormously. A 2025 chapter on eco-friendly AI training highlights strategies like model pruning, energy-efficient hardware, and green data centers as proven ways to cut emissions [4]. A 2022 paper from Google researchers predicts that if the whole machine learning field adopts best practices—such as using more efficient hardware and optimizing training schedules—total carbon emissions from training could actually decline by 2030, despite growing demand [8]. This shows that measurement is not just about accounting; it can drive real reductions.
Where is the evidence still unclear?
The biggest gap is the lack of consistent reporting on chip manufacturing and hardware lifecycle emissions. While several papers call for including 'embodied emissions' from manufacturing [3], none of the nine studies here provide specific data on how much chip fabrication contributes relative to training or inference. This is a critical missing piece because manufacturing advanced AI chips is energy-intensive and produces its own carbon footprint.
Another unresolved issue is how to fairly compare different models and use cases. A 2024 review of neural transpilers found that out of 17 primary studies, only nine reported training time and device details, and even fewer reported actual emissions [5]. This patchy reporting makes it hard to know whether, for example, a small model trained for a long time is better or worse than a large model trained briefly. The authors urge 'more detailed reporting on emissions and a better understanding of the carbon footprint associated with training time' [5]. Until reporting standards become universal, any single number—whether from training, inference, or manufacturing—should be treated as provisional.
About These Sources
This answer is built on 8 peer-reviewed studies — published from 2022 to 2026, 6 from 2024 or later, 2 in Q1 journals, collectively cited 370 times — selected as the most relevant from 9 studies that passed quality screening, drawn from 68 papers retrieved from a database of over 500 million.
Sources used in this answer
Assessing the carbon footprint of language models: Towards sustainability in AI
This 2026 study compares training and inference emissions for small language models (TinyLlama, nanoGPT) and finds that existing benchmarks rarely report energy consumption, creating a significant information gap.
When machine learning models retire, decay, or become obsolete: A review on algorithms, software, and hardware
This 2025 review introduces a six-dimensional framework for sustainable AI (including grid-optimized scheduling, resource-efficient architectures, and energy-aware algorithms) and calls for standardized methods to measure AI energy consumption across the lifecycle.
Toward Sustainable AI: A Scoping Review of Carbon Footprint and Environmental Impacts Across Training and Inference Stages
This 2026 scoping review finds that inference-stage emissions have received far less attention than training, and proposes life-cycle monitoring that includes embodied emissions from manufacturing.
Reducing the Carbon Footprint in Machine Learning With Eco-Friendly AI Training
This 2025 chapter reviews strategies to reduce AI's carbon footprint, including model optimization, pruning, energy-efficient hardware, and green data centers, with case studies from Google, OpenAI, and IBM.
Assessing Carbon Footprint: Understanding the Environmental Costs of Training Devices
This 2024 review of neural transpilers found that out of 17 primary studies, only nine reported training time and device details, and even fewer reported actual emissions, highlighting the need for better reporting.
The Energy and Environmental Footprint of AI
This 2026 article compares four AI environmental footprint calculators and finds that estimates for the same query vary by more than 50-fold, showing the need for standardized disclosure.
Reducing the Carbon Impact of Generative AI Inference (today and in 2035)
This 2023 study models ChatGPT inference workloads and shows that cumulative energy and carbon impacts of inference can exceed those of training, especially as user numbers grow.
The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink
This 2022 paper from Google researchers shows that adopting four best practices (e.g., efficient hardware, optimized training) could cause total ML training carbon emissions to decline by 2030 despite growing demand.
