What you get instead of the publisher's file
Numbers checked against the live data on
Every emission factor here started as a publisher's own file: a spreadsheet from ADEME, a workbook from DESNZ, a set of EPD documents from a construction database. Those files do not agree with each other. They name places differently, print units differently, apply different warming potentials, and mean different things by a total.
This section is what happens in between. What open-climate.ai standardised is the list of changes made on the way in, the model is the shape the data lands in, and the methodology pages explain the concepts a factor carries once it is there.
Why not just use the publisher's file#
You can, and for a single number from a single publisher that is often the right thing to do. Every library page links the publisher's own source, and the Open data page hands you the standardised copy of it.
Three things get hard as soon as you have more than one publisher:
- Comparing. One file says
kg CO2e/kWh, another sayskg CO2e/MWh, a third prices the same fuel per litre. One saysFrance, another saysFR, a third saysEurope. Until units and regions are coded the same way, a comparison is a guess. - Adding up. Two factors on different GWP reports or different system boundaries are not on the same scale, and nothing in the source files says so. Adding them quietly produces a wrong total.
- Defending a number. An auditor asks where a value came from. A cell reference into a spreadsheet you downloaded last year is not an answer if the publisher has since reissued the file.
Nothing is lost by taking the standardised copy. The publisher's own row is kept beside every value, so the original is always one query away: see Trace a served value back to its source line.
What open-climate.ai standardised#
One theme per row below. Each is a change made to the publisher's data so it can be searched and compared with every other library, and each links to the page that explains it in full.
| Theme | What it means | Learn more |
|---|---|---|
| One category axis | Every factor value is sorted into one shared category. The publisher's own classification is kept beside it. | Category |
| Coded regions | Every place is an ISO or registry code. Values that take the publisher's default region say so, rather than assume it. | Region from the publisher's default |
| Canonical units | Every value has one canonical unit, grouped in unit families, with the publisher's raw unit kept. | Units |
| Warming potentials | Values recalculated on AR6, already on AR6, left on the publisher's GWP, or the same on every GWP. | GWP |
| Declared boundaries | Each factor carries a boundary code, and the terms that code establishes say whether they are in, out or varies. | System boundaries |
| Parts and parents | Parts are grouped under their parent factors. Use a part on its own, or the parent when you want the whole. | Parents, children, double counting |
| Summed totals | Totals summed from the parts the publisher printed. | Derived total |
| Spend factors | Factors per amount spent, each with the year its prices are in. | Spend-based vs activity-based |
| Electricity mix | An electricity factor states its grid mix where the publisher accounts for one, with one default value, market-based first. A spend factor priced per euro has no accounting mix to state. | Market-based vs location-based electricity |
| Data quality | The publisher's quality ratings, put on one comparable tier. | Data quality |
| EPDs and suppliers | Values from an EPD, and values that name a supplier. | Supplier |
| Languages | The share of factor values named in each language. | How factors are cleaned |
| Traceable to the source | The publisher's rows kept on file, and the rows not imported, each with a reason. | Nothing was silently dropped |
| Stable addresses | Factor addresses never break: an old address redirects to the current one. | A published slug always resolves |
Not every theme applies to every library. A library with no spend factors has nothing under Spend factors.
The same list, counted per library#
Each library page carries this list as tiles under What open-climate.ai standardised, with the count that theme covers in that library, in factor values, parts or parent factors. A tile counts the library's default slice: one value per factor, on the default indicator and the latest year. Each tile also links to the audit checks behind its claim, so the theme's promise and the tests of that promise sit next to each other.
How factors are cleaned is the same story told transformation by transformation, with the column that lets you undo each one.
What is promised, and what is checked#
Standardising data is only useful if the result is predictable. Two pages say what you can count on:
- Guarantees you can rely on states the promises the model keeps, most of them with a query that tests them against the live data.
- See what was not imported, and why lists what a library left out. An absence always has a recorded reason; a withheld value is unknown, not zero.
What is not changed#
- The publisher's numbers. A value is recalculated only where the change is declared, such as a unit conversion or a GWP restatement, and both the input and the conversion are kept.
- The publisher's terms. The licence travels with every row, and a library whose terms forbid redistribution stays searchable with its values withheld.
- The publisher's own labels. The source classification, the raw unit and the original text stay on the row beside the standardised fields.
In this section
- The model at a glanceWhat a row of factors_flat holds, read as the factor page shows it (app mode) and as the view serves it (SQL mode), with the five ideas behind every row.
- How factors are cleanedEvery transformation applied to publisher data on the way into open_ef, from units and regions to GWP and gases, and how you undo each one from the row.
- Guarantees you can rely onTen promises the open_ef model keeps, from slugs that always resolve to licences on every row, most with a query that checks them against the live data.
- Parents, children, double countingHow open_ef links a total to its parts with parent_id, parent_relation and root_slug, and which filters keep an aggregate from counting emissions twice.
- Choosing a value: the axesWhat the default row hides, the axes a factor's values vary on, and how to pick another GWP set, price basis, grid mix or year instead.
Methodology
- Emission factor methodology
- Market-based vs location-based electricity
- Estimated electricity legs: losses and the fuel chain
- Estimated electricity legs: parameters and sources
- GWP values and IPCC report versions (AR4, AR5, AR6)
- System boundaries: well-to-wheel, cradle-to-gate, EN 15804
- Spend-based vs activity-based emission factors
- Units and unit conversions in emission factors
- Reference year, applicable years and validity dates
- Data quality ratings and the normalised tier
- Licences and redistribution of emission factors
- Biogenic carbon, land use and the CO2e basis
- The data terms and open-climate.ai's own licence
- Waste treatment routes and the elected default
- Regions