Abstract
Dividing an economic outcome by energy is easy. Deciding what the quotient means is harder. Existing work already reports gross domestic product per unit of energy, studies useful work and exergy conversion, constructs energy efficiency indicators, and joins physical with monetary accounts. This paper does not claim those ideas as new. It gives a formal account of the declarations needed before a value-per-joule statistic can support comparison. The central object is a measurement tuple : system boundary, counterfactual, value functional, time horizon, energy convention, attribution rule, and uncertainty model. Conditional on this tuple, a system has The quotient has a clear unit and a reproducible interpretation only when every component is declared. Several results delimit what the statistic can do. A common positive rescaling of a value unit preserves rankings but changes cardinal magnitudes. Heterogeneous outcome vectors that do not dominate one another admit opposing rankings under admissible positive weights. Enlarging a physical boundary, changing a counterfactual, altering the time horizon, or reallocating a shared burden can also reverse a ranking. These are not defects that better arithmetic removes. They identify the normative and accounting choices on which the result depends. The paper supplies a comparability gate, an uncertainty protocol, worked counterexamples, and a standard-library Python reference implementation. The proposed framework is a discipline for making conditional ratios legible, not an energy theory of value and not a universal ranking of industrial, computational, monetary, or cognitive systems.
The quotient and the question
Energy ratios have a reassuring appearance. A numerator is divided by joules, and the result seems to say how much an activity produces from a scarce physical input. The arithmetic can be correct while the comparison is false. One study may count electricity at a device. Another may count electricity at a facility, including cooling and power conversion. A third may add manufacturing energy. Their denominators are all measured in joules, but they describe different systems. The same problem appears above the line. Revenue, consumer surplus, avoided loss, completed tasks, and useful mechanical work are different quantities. Calling each one “value” does not make their units match.
The difficulty is old. Patterson’s survey of energy efficiency indicators distinguishes thermodynamic, physical, economic-thermodynamic, and economic measures, each suited to a different question [1]. Ayres and Warr argue that converted useful work, rather than raw energy alone, matters for historical production [2]. Warr and coauthors estimate useful work over a century in four economies [3], and Warr and Ayres later connect useful work with information in a growth model [4]. National statistical practice already publishes energy intensity and integrated environmental-economic accounts [8, 9]. Thus neither energy-normalized output nor the economic importance of useful work begins here.
This paper asks a narrower question: what must be fixed before a ratio of economic value to energy is interpretable, reproducible, and comparable? The answer is a seven-part measurement tuple. The tuple separates physical accounting from valuation and causal attribution. It also makes non-invariance visible. A reported ranking is conditional on the tuple unless an analyst proves robustness over a declared set of alternatives.
Three distinctions organize the argument.
Physical conversion efficiency compares commensurate physical inputs and outputs. Economic value per joule places a declared value functional in the numerator. The second does not follow from the first.
A ratio can be well defined for one study without being comparable to a ratio from another study. Reproducibility and cross-study comparability are separate gates.
Dependence on a value functional is not mere measurement noise. When outcomes are heterogeneous, a scalar ranking contains a weighting judgment.
The paper contributes four things. First, it defines the tuple and a typed ratio. Second, it proves constructive rank-reversal and non-invariance results. Third, it turns those results into a measurement protocol. Fourth, it maps the formal objects to executable, unit-checked examples. The claims are modest by design. A careful conditional statistic is more useful than a sweeping quotient whose numerator and denominator move whenever the application changes.
Primitives and scope
Definition 1 (System). A system is an intervention, technology, organization, or bounded process for which an analyst defines a functional unit, an outcome record, and an energy inventory over a time interval.
The word “system” does not imply a closed thermodynamic system. Most economic applications are open systems that exchange matter, energy, information, and money with their surroundings. A boundary is an accounting choice around part of that open system.
Definition 2 (Outcome vector). For system , let denote a vector of consequences measured in their native units. Components may include tonnes moved, service hours, correctly completed tasks, dollars of producer surplus, or avoided expected loss. Components with different units remain distinct until a value functional maps them to a scalar.
Definition 3 (Value functional). A value functional maps an outcome record to a declared scalar value space . The codomain may be real currency of a stated price year, a money-metric welfare measure, or a named index. The functional includes its population, standing, discounting, distributional treatment, and treatment of external effects.
A functional contains more than a unit label. “2026 dollars” still leaves open whose willingness to pay counts, whether costs are netted, which externalities enter, and how future values are discounted. Two studies can share a currency and use different functionals.
Definition 4 (Counterfactual increment). Let denote an explicit alternative to system . The incremental value is or the corresponding unit-level causal contrast when outcomes are stochastic.
Definition 5 (Energy inventory). For boundary , energy convention , system , and horizon , let be a non-overlapping energy total in joules. The superscript records the convention, such as device operational, system operational, facility operational, lifecycle, primary, final, or marginal-grid energy.
Strict positivity excludes division by zero. A net energy exporter can still be studied, but its gross inputs and exported energy must be recorded separately. A signed “net joule” denominator would change the interpretation and is outside the present definition.
Definition 6 (Measurement tuple). A value-per-joule measurement tuple is where:
is the physical and organizational system boundary;
is the counterfactual;
is the value functional;
is the time horizon and discounting interval;
is the energy convention;
is the causal and joint-production attribution rule;
is the uncertainty model.
Definition 7 (Value per joule). For a fixed tuple and positive energy inventory, If energy is itself stochastic, the analyst must declare whether the estimand uses expected numerator over expected denominator, an expected unit-level ratio, or another functional. These are generally different.
Remark 8 (Dimensional type). If returns 2026 US dollars and energy is in joules, then has unit 2026 US dollars per joule. If returns successful task equivalents, the unit is successful task equivalents per joule. The two ratios cannot be compared without an additional valuation map.
Why seven fields
Each tuple component blocks a distinct ambiguity. Boundary says what equipment, infrastructure, upstream process, and downstream process belong to the study. Counterfactual determines incrementality. Functional defines the numerator. Horizon controls when benefits and energy enter. Convention distinguishes stages of energy accounting. Attribution allocates shared causes, assets, and burdens. Uncertainty model states what is random, how dependence is represented, and what interval accompanies the point estimate.
None can be inferred reliably from the others. A lifecycle boundary does not select a counterfactual. A randomized counterfactual does not decide how to allocate a shared data center. A price-year label does not disclose a welfare functional. The tuple is intentionally redundant with a good study protocol: its purpose is to make omissions visible.
Well-posedness and the comparability gate
Definition 9 (Internally well posed). A reported is internally well posed if:
every tuple component is declared;
numerator and denominator refer to the same functional unit and horizon;
the energy inventory is positive and contains no overlapping flows;
the value unit and energy unit are stated;
the uncertainty estimand matches the reported estimator.
Definition 10 (Direct comparability). Two results are directly comparable when their functional units match and their tuples are equal, except for representational transformations that have been proved order preserving. A documented bridge may replace exact equality for an energy convention or value unit when the bridge is applied to both results.
Proposition 11 (Exact tuple matching is an equivalence relation). On the set of internally well-posed results, the relation “has the same functional unit and measurement tuple” is reflexive, symmetric, and transitive.
Proof. Equality of the functional unit and each tuple component is reflexive, symmetric, and transitive. Their finite conjunction has the same properties. ◻
PROVED. The result is elementary, but it matters operationally. It lets a registry partition results into comparison classes before anyone sorts their numerical values.
A practical gate
Before ranking two systems, an analyst should answer the following questions.
Do the functional units describe the same service?
Do boundaries include compatible physical components?
Are energy stages identical or connected by a documented bridge?
Do counterfactuals describe the same alternative?
Does use the same population, prices, weights, and discount rule?
Do horizons match?
Are idle energy, failures, shared infrastructure, and embodied burdens allocated under compatible rules?
Are uncertainty intervals based on compatible estimands?
If one answer is no, the comparison is not automatically useless. It becomes a sensitivity exercise or an OBSTRUCTED direct comparison. The analyst can report each result in its own class, derive a bridge, or show a range over plausible tuples. What is not justified is a point ranking that suppresses the difference.
| Field | Result A | Result B |
|---|---|---|
| Boundary | accelerator device | facility meter including cooling |
| Energy convention | final electricity | primary energy equivalent |
| Counterfactual | no automation | older automated system |
| Value functional | private operating profit | social surplus net of emissions |
| Horizon | one benchmark run | five-year lifecycle |
| Attribution | marginal energy | average shared-infrastructure allocation |
| Uncertainty | conditional bootstrap | engineering tolerance only |
Scale transformations and what invariance means
The word “invariant” needs care. A numerical ratio in dollars per joule changes when dollars become cents. A ranking need not change. Cardinal invariance and ordinal invariance are different properties.
Theorem 12 (Common positive value-unit rescaling). Fix . Let for one , applied to every system. Then For any systems ,
Proof. Linearity of subtraction and expectation gives . The energy denominator is unchanged. Division yields the first equality. Multiplication by a common positive scalar preserves strict and weak order. ◻
PROVED. Converting dollars to cents changes every magnitude by 100 but preserves a ranking. This is the appropriate currency-rescaling invariance test. Using different exchange-rate or purchasing-power adjustments for different systems is not the transformation in the theorem.
Proposition 13 (Affine translations cancel only under matched contrasts). If , with common and the same applied to both the intervention and its counterfactual, then . If the translation differs between intervention and baseline, cancellation fails.
Proof. Under a common translation, Unequal translations leave their difference in the numerator. ◻
The proposition explains why the counterfactual is part of unit handling. Incremental value is invariant to a common origin shift, but a gross revenue ratio need not represent incremental value at all.
Non-invariance and rank reversal
Valuation weights
Suppose a system and its counterfactual produce heterogeneous outcomes. Write the componentwise increment as . For nonnegative weights , consider the linear functional . Define the incremental-outcome-per-energy vector where division acts componentwise.
Theorem 14 (Scalarization rank reversal). Let systems and have positive energy denominators. If neither nor weakly dominates the other, then there exist nonnegative, nonzero weight vectors and such that
Proof. Failure of weak dominance in both directions gives coordinates and with and . Choose as the -th coordinate vector and as the -th coordinate vector. Both are nonnegative and nonzero. Then , while . ◻
PROVED. The theorem does not say every small change of weights reverses a ranking. It says a universal scalar order is unavailable for non-dominating vectors unless the admissible valuation class is restricted. The restriction may be ethically or institutionally justified, but it must be published.
Corollary 15 (Robust ranking under componentwise dominance). If componentwise, then for every nonnegative . If at least one strictly better component receives positive weight, the ranking is strict.
Proof. Every component of is nonnegative, so its inner product with a nonnegative vector is nonnegative. Strictness follows when one positive difference has positive weight. ◻
Dominance is the strongest defensible cross-weight conclusion. When it fails, a Pareto frontier communicates more than a single league table.
System boundaries
Theorem 16 (Boundary expansion can reverse a ranking). Let two systems have positive values and operational energies . Suppose If an admissible expanded boundary assigns additional energy to system and none to , then the ranking reverses whenever
Proof. The expanded-boundary ranking favors exactly when All denominators and values are positive. Cross multiplication and rearrangement give the stated threshold. ◻
PROVED, conditional on the accounting admissibility of the boundary expansion. The theorem is not permission to add arbitrary energy to a disfavored system. The added burden must belong to the declared expanded boundary under a consistent allocation rule. Examples include manufacturing energy for dedicated hardware or a facility load omitted from a device measurement.
Remark 17 (Symmetric expansions). Both systems may receive added burdens. A reversal occurs when No qualitative result follows from the word “lifecycle” alone. The quantities and , their lifetimes, and their allocation across functional units determine the comparison.
Counterfactuals
Proposition 18 (Gross outcomes do not identify incremental rankings). On an unrestricted scalar outcome space, let observed outcomes satisfy . For any desired strict ordering of incremental values, there exist counterfactual outcomes and that produce it through and .
Proof. Choose any target increments with the desired ordering, then set and . The observed outcomes remain fixed while the increments equal the targets. ◻
PROVED as a non-identification construction. Empirical counterfactuals cannot be chosen after seeing the desired answer. The construction shows why observed revenue per joule alone does not identify incremental value per joule. Design, institutional knowledge, or defensible bounds must restrict the counterfactual class.
Time horizons
Proposition 19 (Horizon reversal). There exist systems with equal energy such that has higher value per joule at horizon , while has higher value per joule at a longer horizon .
Proof. Let both systems use one joule at time zero. Let produce value 10 at time one and no later value. Let produce value 6 at time one and 8 at time two. At horizon one, the ratios are 10 and 6. At horizon two, with no discounting, they are 10 and 14. Positive discounting preserves a reversal for a nonempty range of discount factors. ◻
PROVED. Delayed maintenance costs, learning effects, degradation, and asset life can generate the same structure. A horizon should therefore be chosen from the decision problem, not from whichever endpoint favors a technology.
No universal scalar order
Theorem 21 (Tuple-relative ordering). Consider a domain that contains at least two systems with non-dominating outcome-per-energy vectors and permits the coordinate valuations used in 14. No complete scalar ranking of that domain is invariant to all admissible value functionals.
Proof. Assume a complete invariant scalar ranking exists. For the two non-dominating systems, completeness ranks one weakly above the other or treats them as tied. 14 supplies one admissible valuation that ranks the first strictly above the second and another that ranks it strictly below. Either strict ranking contradicts a fixed tie, and one contradicts either fixed strict order. Thus the assumed invariant ranking does not exist. ◻
PROVED under the stated valuation class. The theorem is deliberately weaker than a social-choice impossibility theorem. It needs no claim about collective preference aggregation. It only records a fact about scalarizing heterogeneous outcomes.
Worked examples
Value weights
Consider two stylized systems that each use 100 joules. Relative to a zero-vector counterfactual, system A adds nine verified task completions and two units of avoided loss. System B adds three verified task completions and eight units of avoided loss. Their incremental-outcome-per-energy vectors are If the functional assigns one index point to each completed task and zero to avoided loss, A ranks first. If it assigns one point to each avoided-loss unit and zero to task count, B ranks first. Neither vector dominates.
| System | Energy (J) | Tasks | Avoided loss | Preferred by |
|---|---|---|---|---|
| A | 100 | 9 | 2 | task-only functional |
| B | 100 | 3 | 8 | loss-only functional |
The correct response is not to average the dimensions by default. One can publish both components, state a reason for weights, and show a sensitivity region. For linear weights , A ranks first when or . B ranks first when , and they tie when the weights are equal. The threshold is transparent.
Operational and lifecycle boundaries
Now give both systems the same realized value of ten currency units. Under an operational boundary, A uses 40 joules and B uses 50. Their ratios are 0.25 and 0.20 currency units per joule, so A ranks first. An expanded boundary assigns 100 joules of embodied energy to A and 20 to B. The lifecycle denominators become 140 and 70. The ratios become approximately 0.0714 and 0.1429, so B ranks first.
| System | Value | Operational E | Embodied E | Operational VPJ | Lifecycle VPJ |
|---|---|---|---|---|---|
| A | 10 | 40 | 100 | 0.2500 | 0.0714 |
| B | 10 | 50 | 20 | 0.2000 | 0.1429 |
This example is COMPUTATIONAL. It verifies the construction in 16; it is not a lifecycle assessment of a real technology. An empirical study would need inventories, lifetimes, utilization, allocation, geography, and uncertainty.
Physical efficiency versus realized value
Suppose heat system H converts 90 percent of final electricity into useful heat, while system L converts 70 percent. On a matched physical functional unit and boundary, H is more physically efficient. Now suppose H serves a process that often produces inventory with no buyer, while L serves a temperature-sensitive repair whose completion avoids a costly outage. L can have higher expected realized value per joule despite lower heat-delivery efficiency.
There is no paradox. Physical efficiency answers how much useful heat leaves the device for a given physical input. Economic VPJ also depends on the use, counterfactual, prices or welfare weights, reliability, and attribution. The physical result remains true within its domain. It is one input to the economic account, not a complete valuation.
The danger of a gross ratio
Two sites report annual revenue and electricity: The first site appears better. But suppose the intervention studied at A replaced a process that already earned 11 currency/MJ, while B replaced one that earned 3 currency/MJ, on matched energy bases. The incremental comparisons are one and six currency/MJ. Gross output and incremental contribution answer different questions.
The numerical subtraction alone does not establish causality. The baselines would need evidence. A randomized rollout, credible natural experiment, engineering model, or partial-identification bound could support them. The example’s role is to make the estimand distinction concrete.
Uncertainty and ratio estimands
Three different ratios
Let and be random. Analysts may encounter: The third is a sample ratio estimator for the first under common sampling conditions. The second is the mean of unit-level ratios. In general, . If high-value observations also use more energy, an unweighted mean of unit ratios can answer a particularly different question from an aggregate ratio. In the notation of 7, is and is ; this subsection relaxes the definition’s default treatment of the denominator as fixed.
Proposition 22 (Expectation does not commute with division). There exist positive random energies and values such that
Proof. Let always and let equal 1 or 2 with equal probability. Then ◻
PROVED. The uncertainty component must name the estimand before it names a confidence-interval algorithm.
Sources of uncertainty
A useful uncertainty ledger separates:
measurement error in meters, prices, and outcome records;
sampling variation across tasks, days, sites, or users;
model uncertainty in counterfactual and causal assumptions;
scenario uncertainty in future demand, lifetime, and discounting;
normative uncertainty over weights and social standing;
boundary uncertainty over omitted or shared processes.
The first two may yield to repeated measurement and conventional intervals. The last four often require sensitivity analysis, scenario sets, or bounds. A narrow bootstrap around one selected model does not represent uncertainty about the model or value functional. Saltelli and coauthors provide a general framework for global sensitivity analysis when uncertain inputs interact [19].
Recommended reporting
For an empirical ratio, report the numerator and denominator separately before the quotient. Include their covariance when both are estimated from the same sample. If the denominator is noisy or can approach zero, ordinary symmetric intervals can fail. The present framework requires positive energy and recommends Fieller-type or bootstrap methods only after checking their assumptions [20]. When causal identification is weak, report bounds on and propagate those bounds through the positive denominator.
Sensitivity to tuple choices is not summarized adequately by one standard error. A robust report contains:
a statistical interval conditional on the chosen tuple;
a scenario range over defensible boundaries and horizons;
a weight region or Pareto statement for heterogeneous outcomes;
an identification statement for the counterfactual;
a list of unresolved exclusions.
Measurement protocol
Step 1: state the decision and functional unit
Start with the decision the metric will inform. “Which server is more efficient?” is incomplete. A functional unit might be one thousand requests meeting a named accuracy and latency threshold, one tonne-kilometre delivered, or one year of a settlement service under a named threat model. The unit should describe a comparable service, not an easy intermediate count.
Record the population and deployment context. A benchmark mixture, geographic site, and service-level constraint can all change outcomes and energy. If the functional unit cannot be matched, stop before forming a comparative ratio.
Step 2: draw the physical boundary
List included equipment and processes. Mark the meter location. Record idle energy, warm-up, failures, retries, power conversion, cooling, networking, and shared infrastructure as included, excluded, or separately allocated. For a lifecycle study, add manufacturing, construction, maintenance, and end-of-life processes without overlapping the operational inventory.
Choose one energy convention. Device operational energy is appropriate for a low-level engineering question. Wall energy fits many deployed computing comparisons. Facility operational energy adds site overhead. Lifecycle energy adds non-overlapping embodied burdens. Primary and final energy require a bridge through conversion losses. Marginal-grid energy is indexed by location and time. These conventions should appear as enum-like labels in data, not as prose buried in a methods appendix.
Step 3: predeclare the counterfactual
Name what would happen without the system. “No energy use” is rarely a credible baseline when another process would supply the service. Specify whether the alternative is no service, manual work, an incumbent technology, delayed action, or a portfolio response.
Then state the identification strategy. Random assignment can identify an average effect for the assigned population under its assumptions. Observational designs need exchangeability, timing, exclusion, or structural assumptions. Engineering counterfactuals need validation. If none is credible, an OPEN or partially identified numerator is more honest than a precise gross ratio.
Step 4: define the value functional
Write as an auditable calculation. For a monetary functional, state:
currency, price year, and conversion method;
whether the measure is revenue, profit, surplus, or cost saving;
whose costs and benefits count;
treatment of taxes, transfers, externalities, and distribution;
discount rate and terminal value;
treatment of risk and catastrophic loss.
For a nonmonetary index, state component units, normalization, weights, and aggregation rule. Publish the component vector. A composite should be a view of the data, not its only surviving form.
Step 5: set horizon and attribution
The horizon should include relevant benefits, operating energy, degradation, maintenance, and replacement. If horizons differ across systems, annualize or discount with a documented bridge. Do not count a full asset benefit against one day of energy or a full manufacturing burden against one unusually small batch.
For joint production, identify dedicated and shared flows. Prefer causal or measured marginal allocation when available. Otherwise show results under more than one defensible rule. Capacity share, run time, revenue share, mass, and equal allocation answer different questions. The allocation rule belongs in the tuple because 20 shows it can decide the ranking.
Step 6: specify uncertainty before estimation
List random variables, dependence, sampling unit, missingness, and estimator. Predeclare whether the target is a ratio of expectations or an expectation of ratios. Record meter accuracy and calibration, data exclusions, and failed observations. Run sensitivity over boundary, baseline, weights, horizon, and allocation, as well as sampling noise.
Step 7: apply the comparability gate
Compare tuples field by field. If a bridge is used, publish the equation, source data, and uncertainty. Report a scalar ranking only inside a comparison class. Outside it, report a vector, Pareto frontier, scenario range, or OBSTRUCTED comparison.
| Field | Required content |
|---|---|
| result identifier | immutable study and scenario identifier |
| functional unit | quantified service and quality threshold |
| boundary | included and excluded components |
| counterfactual | alternative and identification design |
| value functional | formula, unit, standing, prices, weights |
| horizon | start, end, discounting, lifetime treatment |
| energy convention | named stage and meter or inventory source |
| attribution rule | joint outputs and shared energy allocation |
| uncertainty model | estimand, random variables, intervals, scenarios |
| numerator | estimate, unit, interval, provenance |
| denominator | estimate, joules, interval, provenance |
| status | proved, conditional, computational, obstructed, or open |
Interpretation and policy use
What a high ratio can mean
Inside a matched tuple, a higher ratio means more expected declared incremental value per included joule. That sentence is intentionally repetitive. Remove “declared,” “incremental,” or “included,” and the claim grows beyond the measurement.
A high ratio can result from a larger causal benefit, a smaller energy inventory, or both. The decomposition matters for policy. A low-energy, low-value activity and a high-energy, high-value activity can have the same ratio but different scale, risk, and constraints. Ratios should therefore be reported beside absolute energy and absolute value.
Average and marginal questions
An average lifecycle VPJ is not the marginal value of the next joule. The latter depends on time, location, constraints, and displacement. Average metrics may support benchmarking or historical accounting. Marginal allocation requires a response function and an opportunity cost. Paper 6 of this program develops that optimization problem.
The distinction also blocks a policy error. A sector with high average value per joule need not be the best recipient of one additional joule if it is at capacity, has diminishing returns, or cannot use energy at the available node and time. Conversely, a low average ratio does not prove that every marginal use is low value.
Implementation and executable evidence
The companion module src/joule_standard/foundations.py uses only the Python standard library. Decimal arithmetic rejects binary floating-point input at the public quantity boundary. Energy is stored in joules and must be strictly positive. Kilowatt-hours convert by the exact identity
The MeasurementTuple data class has fields for the seven declarations in 6. ValueFunctional carries a name and output unit. IncrementalValue carries an exact amount and unit. The value_per_joule function rejects a numerator whose unit differs from the declared functional. A ValuePerJoule comparison rejects unequal tuples or value units.
Incremental outcome vectors, with a zero vector as the counterfactual in the published fixtures, and nonnegative linear valuations implement 14. The fixtures reproduce the weight reversal in 2 and the boundary reversal in 3. Tests also check:
exact joule and kilowatt-hour conversion;
presence of all seven tuple fields;
positive-energy and unit mismatch rejection;
refusal to compare different boundaries;
common currency-rescaling order invariance;
strict valuation-weight rank reversal;
strict boundary rank reversal;
missing outcome-dimension rejection.
These tests are COMPUTATIONAL. The algebraic theorems have separate proofs in the paper. Passing a finite test suite does not prove a universal statement, and the implementation does not estimate empirical values on its own.
Limitations and nonclaims
Limitation 23 (No natural welfare unit). The framework requires ; it does not derive a unique from thermodynamics. Value is not conserved like energy. Money, utility, capability, rights, security, and task success do not become one physical dimension because each is divided by joules.
Limitation 24 (No energy theory of value). Energy is necessary for material production and computation, but energy input alone does not determine exchange value or welfare. Scarcity, preferences, institutions, knowledge, location, timing, complementary inputs, and distribution matter. This paper does not revive a caloric theory of price.
Limitation 25 (No universal sector ranking). The rank-reversal results rule out the paper’s own use as a universal league table for artificial intelligence, cryptographic settlement, manufacturing, transport, health, or other sectors. Cross-sector comparison requires a common functional unit or an explicit welfare functional and matched boundaries.
Limitation 26 (No causal effect from an observational ratio). Observed value divided by observed energy is descriptive. It becomes an incremental causal ratio only under a design or identification argument that supports the counterfactual numerator. Adding more decimal places does not fix confounding.
Limitation 27 (No automatic lifecycle superiority). A broader boundary can be relevant, but breadth is not accuracy by itself. Lifecycle inventories contain allocation, lifetime, geography, and data-quality choices. Operational and lifecycle figures should be labeled and used for their respective questions.
Limitation 28 (No sufficiency of the ratio). Even a well-posed VPJ statistic omits absolute scale, budget constraints, nonconvexities, tail risk, rights, feasibility, and distribution unless and the decision model include them. A ratio is one statistic, not a complete social decision procedure.
Empirical scope left open
This foundation paper does not estimate a national energy-productivity series, an AI benchmark, or proof-of-work assurance. Those applications need domain specific functional units and data. Their empirical status remains OPEN here. The absence is deliberate: a synthetic example should not be mistaken for a measured sector claim.
Research implications
The tuple provides a common interface for later work without forcing common units. A typed transduction model can track energy, computation, immediate output, verified outcome, adoption, and realized value as separate stages. Quality-adjusted machine intelligence can use task success per wall joule while stopping short of realized business value. Proof-of-work analysis can measure scenario-conditioned settlement assurance without calling energy itself “trust.” Marginal allocation can operate on response curves rather than average ratios. National accounts can publish vectors and bridge tables before any composite.
Several research questions remain.
Which tuple components can be standardized across domains, and which must remain application specific?
How should partial identification of the numerator interact with inventory uncertainty in the denominator?
What dominance or robust-order regions survive over a publicly defensible class of welfare weights?
How should shared digital and physical infrastructure be allocated when marginal use, reservation capacity, and causal responsibility disagree?
Which provenance format lets auditors reproduce every bridge without exposing confidential unit-level data?
The most useful near-term output is not one global number. It is a registry of typed results whose comparison classes, bridges, and sensitivity regions are machine readable.
Conclusion
Value per joule is meaningful only after the analyst says whose value, relative to what alternative, over which period, inside which boundary, under which energy convention, with what attribution rule, and under what uncertainty model. The tuple records those choices.
Once the tuple is fixed, the quotient is ordinary and useful. A common positive currency rescaling preserves order. Matched tuples define clear comparison classes. Numerator and denominator can be audited separately. The statistic can support engineering, operating, investment, or policy analysis within its declared scope.
The same formalism explains the limit. Non-dominating outcomes reverse under admissible weights. Boundaries, counterfactuals, horizons, and shared-burden allocations can reverse rankings too. These dependencies are not an excuse to abandon measurement. They are instructions for honest measurement: publish the tuple, retain the outcome vector, test sensitivity, and decline comparisons that fail the gate.
Proof details and extensions
Weight regions for two outcomes
Let the normalized difference between two systems be , with . Normalize nonnegative weights so . System ranks first when Solving gives The right side lies strictly between zero and one. Thus the admissible simplex contains a nonempty region favoring each system and one tie point. This gives a complete sensitivity diagram for two components.
With , the robust-order region is the intersection of the weight simplex with a half-space: Publishing this polytope is often more informative than publishing one selected weight vector.
Strictly positive weights
14 uses coordinate vectors, which permit zero weights. If the admissible class requires every component to have positive weight, the reversal still holds for sufficiently small perturbations. Let select a coordinate where is better. Define for . Continuity of the inner product preserves the strict inequality for all sufficiently small . Apply the same construction at a coordinate where is better.
Boundary threshold with unequal values
For positive , a burden pair favors when Solving for gives The expression is a sensitivity threshold, not an empirical estimate. A study can place distributions or inventory intervals on and , then report the probability or scenario share on either side.
Discounted horizon reversal
In 19, let the time-two discount factor be . System B’s long-horizon value is . It exceeds A’s value of 10 whenever . Thus the reversal persists for discount rates whose one-period factor exceeds one half. The result does not select a discount rate. It identifies the threshold at which the decision changes.
Ratio decomposition
For two tuples that differ in both numerator and denominator, a logarithmic decomposition can be useful when values remain positive: This identity separates value change from energy change. It does not attribute causality, and it fails when incremental value is zero or negative. Those cases should be reported directly rather than repaired by an arbitrary offset.
Boundary manifest template
An empirical study can use the following template.
Study identifier. Immutable name, version, date, and owner.
Decision. The choice the statistic is meant to inform.
Functional unit. Quantity, quality threshold, location, and service conditions.
Physical system. Diagram or inventory of included processes.
Temporal interval. Measurement window, benefit horizon, lifetime, and discounting.
Geography. Site, grid node or region, and market jurisdiction.
Counterfactual. Alternative process and evidence supporting it.
Energy stage. Device, system, facility, lifecycle, primary, final, or marginal grid.
Meter and inventory. Instrument, calibration, sampling rate, missingness, and data lineage.
Idle and failed work. Inclusion and allocation.
Embodied energy. Included assets, lifetime, utilization, and allocation.
Value functional. Formula, unit, prices, standing, external effects, distribution, and risk.
Attribution. Causal design and joint-production allocation.
Uncertainty. Estimand, stochastic model, scenarios, and sensitivity ranges.
Exclusions. Every known omitted process or consequence.
Comparability worksheet
| Check | Result A | Result B | Disposition |
|---|---|---|---|
| Functional unit | match, bridge, stop | ||
| Boundary | match, scenario | ||
| Counterfactual | re-estimate | ||
| Value functional | weight region | ||
| Horizon | discount bridge | ||
| Energy convention | energy bridge | ||
| Attribution | sensitivity | ||
| Uncertainty estimand | restate |
The disposition “stop” means stop the direct point ranking, not stop the research. A mismatch can motivate a new measurement or a useful sensitivity analysis.
Reproduction record
The reference tests run with Python 3.11 or later and pytest. From the repository root:
python -m pytest tests/test_foundations.py
The code uses no runtime dependency outside the standard library. Tests use pytest as the runner. Exact decimal-like inputs must be integers, strings, or Decimal values; public constructors reject binary floats. This is a small guard against pretending that a display-rounded decimal is exact.
The two rank-reversal fixtures are intentionally transparent. They do not depend on random seeds, hidden data, or fitted parameters. Their purpose is to make the existence proofs executable. Any empirical application should live in a separate module with source provenance and should not replace these fixtures with paper-specific hard-coded acceptance checks.
Claim ledger
| Identifier | Claim | Status |
|---|---|---|
| JS-C001 | VPJ is well defined only relative to a declared tuple and positive, typed denominator. | PROVED |
| JS-C002a | Non-dominating outcome vectors admit valuation-weight reversal. | PROVED |
| JS-C002b | An admissible boundary expansion can reverse a ranking. | PROVED |
| JS-C002c | Gross outcomes do not identify incremental-value rankings. | PROVED |
| JS-C002d | Horizon and shared-burden allocation can reverse rankings. | PROVED |
| JS-COMP-1 | The published Python fixtures reproduce currency scaling and two rank reversals. | COMPUTATIONAL |
| JS-EMP-1 | Any real sector ranking under this framework. | OPEN |
Nonclaim checklist
For clarity, the following statements are outside the paper’s claims:
GDP divided by energy is new.
Useful-work economics is new.
Joules are a currency, utility unit, or conserved measure of value.
Physical conversion efficiency determines economic value.
The framework supplies a politically neutral welfare functional.
Artificial intelligence, Bitcoin, manufacturing, and transport have a natural universal ranking.
An observational value-energy ratio identifies a causal effect.
Lifecycle accounting eliminates allocation choices.
A finite Python test suite proves the general theorems.
A high average ratio proves that the next joule should be allocated to the same system.
References
M. G. Patterson. What is energy efficiency? Concepts, indicators and methodological issues. Energy Policy, 24(5):377 to 390, 1996. https://doi.org/10.1016/0301-4215(96)00017-1.
R. U. Ayres and B. Warr. Accounting for growth: The role of physical work. Structural Change and Economic Dynamics, 16(2):181 to 209, 2005. https://doi.org/10.1016/j.strueco.2003.10.003.
B. Warr, R. U. Ayres, N. Eisenmenger, F. Krausmann, and H. Schandl. Energy use and economic development: A comparative analysis of useful work supply in Austria, Japan, the United Kingdom and the US during 100 years of economic growth. Ecological Economics, 69(10):1904 to 1917, 2010. https://doi.org/10.1016/j.ecolecon.2010.03.021.
B. Warr and R. U. Ayres. Useful work and information as drivers of economic growth. Ecological Economics, 73:93 to 102, 2012. https://doi.org/10.1016/j.ecolecon.2011.09.006.
International Energy Agency. Energy End-uses and Efficiency Indicators Data Explorer. https://www.iea.org/data-and-statistics/data-tools/energy-end-uses-and-efficiency-indicators-data-explorer. Accessed 23 July 2026.
International Organization for Standardization. ISO 14040:2006, Environmental management: Life cycle assessment, principles and framework. https://www.iso.org/standard/37456.html.
International Organization for Standardization. ISO 14044:2006, Environmental management: Life cycle assessment, requirements and guidelines. https://www.iso.org/standard/38498.html.
United Nations, European Commission, Food and Agriculture Organization, International Monetary Fund, OECD, and World Bank. System of Environmental-Economic Accounting 2012: Central Framework. United Nations, 2014. https://unstats.un.org/unsd/envaccounting/seeaRev/SEEA_CF_Final_en.pdf.
United Nations Statistics Division. System of Environmental-Economic Accounting for Energy. https://seea.un.org/content/seea-energy.
M. Fleurbaey. Beyond GDP: The quest for a measure of social welfare. Journal of Economic Literature, 47(4):1029 to 1075, 2009. https://doi.org/10.1257/jel.47.4.1029.
M. Fleurbaey and D. Blanchet. Beyond GDP: Measuring Welfare and Assessing Sustainability. Oxford University Press, 2013. https://doi.org/10.1093/acprof:oso/9780199767199.001.0001.
J. E. Stiglitz, J. P. Fitoussi, and M. Durand. Beyond GDP: Measuring What Counts for Economic and Social Performance. OECD Publishing, 2018. https://doi.org/10.1787/9789264307292-en.
W. E. Diewert. Exact and superlative index numbers. Journal of Econometrics, 4(2):115 to 145, 1976. https://doi.org/10.1016/0304-4076(76)90009-9.
OECD, European Union, and European Commission Joint Research Centre. Handbook on Constructing Composite Indicators: Methodology and User Guide. OECD Publishing, 2008. https://doi.org/10.1787/9789264043466-en.
D. B. Rubin. Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5):688 to 701, 1974. https://doi.org/10.1037/h0037350.
P. W. Holland. Statistics and causal inference. Journal of the American Statistical Association, 81(396):945 to 960, 1986. https://doi.org/10.1080/01621459.1986.10478354.
G. W. Imbens and D. B. Rubin. Causal Inference for Statistics, Social, and Biomedical Sciences. Cambridge University Press, 2015. https://doi.org/10.1017/CBO9781139025751.
C. F. Manski. Identification for Prediction and Decision. Harvard University Press, 2007. https://www.hup.harvard.edu/books/9780674026537.
A. Saltelli, M. Ratto, T. Andres, F. Campolongo, J. Cariboni, D. Gatelli, M. Saisana, and S. Tarantola. Global Sensitivity Analysis: The Primer. Wiley, 2008. https://doi.org/10.1002/9780470725184.
E. C. Fieller. Some problems in interval estimation. Journal of the Royal Statistical Society, Series B, 16(2):175 to 185, 1954. https://doi.org/10.1111/j.2517-6161.1954.tb00159.x.
G. Rebitzer et al. Life cycle assessment, part 1: Framework, goal and scope definition, inventory analysis, and applications. Environment International, 30(5):701 to 720, 2004. https://doi.org/10.1016/j.envint.2003.11.005.
B. W. Ang. The LMDI approach to decomposition analysis: A practical guide. Energy Policy, 33(7):867 to 871, 2005. https://doi.org/10.1016/j.enpol.2003.10.010.
N. Georgescu-Roegen. The Entropy Law and the Economic Process. Harvard University Press, 1971. https://doi.org/10.4159/harvard.9780674281653.
A. Sen. Collective Choice and Social Welfare. Holden-Day, 1970. Expanded edition, Harvard University Press, 2017. https://www.hup.harvard.edu/books/9780674919211.
European Commission, International Monetary Fund, OECD, United Nations, and World Bank. System of National Accounts 2008. https://unstats.un.org/unsd/nationalaccount/docs/SNA2008.pdf.