What you get instead of the publisher's file

Numbers checked against the live data on

Every emission factor here started as a publisher's own file: a spreadsheet from ADEME, a workbook from DESNZ, a set of EPD documents from a construction database. Those files do not agree with each other. They name places differently, print units differently, apply different warming potentials, and mean different things by a total.

This section is what happens in between. What open-climate.ai standardised is the list of changes made on the way in, the model is the shape the data lands in, and the methodology pages explain the concepts a factor carries once it is there.

Why not just use the publisher's file#

You can, and for a single number from a single publisher that is often the right thing to do. Every library page links the publisher's own source, and the Open data page hands you the standardised copy of it.

Three things get hard as soon as you have more than one publisher:

  • Comparing. One file says kg CO2e/kWh, another says kg CO2e/MWh, a third prices the same fuel per litre. One says France, another says FR, a third says Europe. Until units and regions are coded the same way, a comparison is a guess.
  • Adding up. Two factors on different GWP reports or different system boundaries are not on the same scale, and nothing in the source files says so. Adding them quietly produces a wrong total.
  • Defending a number. An auditor asks where a value came from. A cell reference into a spreadsheet you downloaded last year is not an answer if the publisher has since reissued the file.

Nothing is lost by taking the standardised copy. The publisher's own row is kept beside every value, so the original is always one query away: see Trace a served value back to its source line.

What open-climate.ai standardised#

One theme per row below. Each is a change made to the publisher's data so it can be searched and compared with every other library, and each links to the page that explains it in full.

Theme What it means Learn more
One category axis Every factor value is sorted into one shared category. The publisher's own classification is kept beside it. Category
Coded regions Every place is an ISO or registry code. Values that take the publisher's default region say so, rather than assume it. Region from the publisher's default
Canonical units Every value has one canonical unit, grouped in unit families, with the publisher's raw unit kept. Units
Warming potentials Values recalculated on AR6, already on AR6, left on the publisher's GWP, or the same on every GWP. GWP
Declared boundaries Each factor carries a boundary code, and the terms that code establishes say whether they are in, out or varies. System boundaries
Parts and parents Parts are grouped under their parent factors. Use a part on its own, or the parent when you want the whole. Parents, children, double counting
Summed totals Totals summed from the parts the publisher printed. Derived total
Spend factors Factors per amount spent, each with the year its prices are in. Spend-based vs activity-based
Electricity mix An electricity factor states its grid mix where the publisher accounts for one, with one default value, market-based first. A spend factor priced per euro has no accounting mix to state. Market-based vs location-based electricity
Data quality The publisher's quality ratings, put on one comparable tier. Data quality
EPDs and suppliers Values from an EPD, and values that name a supplier. Supplier
Languages The share of factor values named in each language. How factors are cleaned
Traceable to the source The publisher's rows kept on file, and the rows not imported, each with a reason. Nothing was silently dropped
Stable addresses Factor addresses never break: an old address redirects to the current one. A published slug always resolves

Not every theme applies to every library. A library with no spend factors has nothing under Spend factors.

The same list, counted per library#

Each library page carries this list as tiles under What open-climate.ai standardised, with the count that theme covers in that library, in factor values, parts or parent factors. A tile counts the library's default slice: one value per factor, on the default indicator and the latest year. Each tile also links to the audit checks behind its claim, so the theme's promise and the tests of that promise sit next to each other.

How factors are cleaned is the same story told transformation by transformation, with the column that lets you undo each one.

What is promised, and what is checked#

Standardising data is only useful if the result is predictable. Two pages say what you can count on:

What is not changed#

  • The publisher's numbers. A value is recalculated only where the change is declared, such as a unit conversion or a GWP restatement, and both the input and the conversion are kept.
  • The publisher's terms. The licence travels with every row, and a library whose terms forbid redistribution stays searchable with its values withheld.
  • The publisher's own labels. The source classification, the raw unit and the original text stay on the row beside the standardised fields.

In this section

Methodology

Advanced