Gini coefficient
Enable an AI agent to identify, interpret, compare and assess the permitted uses of a Gini coefficient by recording the distribution measured, calculation convention and evidential limitations.
Research draft, second pass
A second pass drafted this model: the structure a model of this thing needs, and what is known about it in the world. The line under this one says how the second half was obtained - researched against sources, or recalled without web access, in which case nothing here was read anywhere and every claim is a lead to verify. Unreviewed either way.
Researched by: Codex + Grok
Purpose and description
Enable an AI agent to identify, interpret, compare and assess the permitted uses of a Gini coefficient by recording the distribution measured, calculation convention and evidential limitations.
The Gini coefficient is a relative, scale-invariant inequality index equal to half the mean absolute pairwise difference of a distribution divided by its mean (equivalently, twice the area between the Lorenz curve and the egalitarian diagonal); for a non-negative variable it is 0 under complete equality and approaches 1 as one unit holds the entire total.
It can be Identify whether a reported Gini refers to distributional inequality and attach its population, resource definition and period.; Convert between declared numerical scales while preserving the source value and calculation convention.; Calculate or reproduce a coefficient from suitable data using an explicit estimator and documented input treatment.; Assess whether a value is mathematically interpretable under its input conditions and normalization.; Compare estimates after checking resource definitions, populations, methods and uncertainty.; Qualify or reject claims that infer poverty, welfare, distributional identity or causal effects from the coefficient alone..
Distinguishing features
Represents inequality or concentration across a distribution of quantities; it does not identify the distribution's average level, total size or poverty rate.
For a nonnegative distribution with positive mean, the conventional coefficient is associated with the area between the Lorenz curve and the equality line and can be expressed using normalized pairwise absolute differences.
A reported value requires a declared scale: a coefficient on a 0-1 scale and an index on a 0-100 scale can express the same estimate.
Decision-tree Gini impurity measures category mixing, while predictive Gini commonly derives from discrimination performance; sharing the name does not make either the distributional coefficient modeled here.
Equal Gini coefficients do not establish identical distributions or Lorenz curves, and a Gini ranking alone does not establish Lorenz dominance.
Scope
+ The quantity whose inequality or concentration is measured and the units across which it is distributed
+ The reported coefficient, numerical scale, formula and normalization convention
+ Population coverage, weighting, data treatment and estimation uncertainty
+ Conditions for comparing coefficients across populations, periods or methods
+ Interpretive limits and decisions about whether an estimate can support a proposed claim
- The underlying household survey, administrative dataset or sampling system as independently managed objects
- Complete income, wealth, consumption or other accounting frameworks
- Poverty, welfare and social justice assessments beyond the inequality information supplied by the coefficient
- Alternative inequality measures such as Theil indices, Atkinson indices and percentile ratios
- Decision-tree Gini impurity and predictive discrimination metrics called Gini
- Policies intended to change a distribution and evaluations of their causal effects
Characteristics
- Reported Gini value
- Dimensionless; record the reported number and whether its scale is 0-1, 0-100 or another declared convention Prevents scale errors and preserves the published estimate without assuming that every input distribution satisfies conventional bounds.
- Distributed quantity
- Income, wealth, consumption, land holdings or another explicitly defined quantity The same numerical coefficient has different substantive meaning when applied to different quantities.
- Resource accounting basis
- For example, gross or disposable income; individual, household, per-capita or equivalised resources; applicable reference period Accounting and allocation choices can change measured inequality before any population change occurs.
- Represented population and counting unit
- Links to the target population, geographic coverage, eligibility rules and units receiving statistical weight Distinguishes inequality among households from inequality among people assigned household resources.
- Estimator and normalization
- Declared formula, weighting procedure, grouped-data approximation and any finite-sample correction Different conventions can yield different values and attainable upper bounds for the same observations.
- Input-domain condition
- Nonnegative with positive mean; all zero; negative observations present; zero or negative mean; unknown Determines whether conventional bounds and interpretations apply and whether the stated formula is defined.
- Statistical uncertainty
- Standard error or interval on a declared Gini scale, with confidence level and estimation method; unavailable when not estimated Supports assessment of whether observed differences exceed sampling uncertainty.
- Comparability assessment
- Comparable for a stated purpose; comparable after documented adjustment; not comparable; unresolved Prevents differences in coverage, accounting or calculation from being mistaken for changes in inequality.
Where this came from
wikidata · CC0 1.0
Drafted structure
Bundle to layer to finding to question, as the second pass will find it: 6 bundles · 11 layers · 18 findings · 28 questions.
Distributional target Identifies what is distributed, among whom and over which reference period.
A Gini value cannot be interpreted until the resource and recipient population are specified.
Resource definition
Captures the substantive quantity and accounting choices underlying the distribution.
Resource and accounting basis
Establishes whether the coefficient measures income, wealth, consumption or another quantity and how that quantity is constructed.
- What quantity is distributed, and which components, taxes, transfers, liabilities or in-kind resources are included where applicable? definition
- Does the quantity represent a stock at a date or a flow over a period, and what date or period applies? boundary
Recipients and coverage
Separates the recorded observation unit from the population represented by the inequality estimate.
Population and resource allocation
Records who is counted and how resources are assigned to those units.
- Which geographic and eligibility boundaries define the population, and which groups are excluded? boundary
- Is inequality calculated across people, households or other units, and are household resources assigned using per-capita division or an equivalence scale? definition
Coefficient convention Identifies the mathematical statistic, reported scale and normalization used.
The Gini name and number alone do not establish which statistic or finite-population convention was applied.
Statistic identity
Distinguishes the distributional coefficient from other statistics sharing its name.
Formula and scale
Connects the reported number to an explicit distributional Gini definition.
- Which formula defines this estimate, and does it measure distributional inequality rather than classification impurity or predictive discrimination? definition
- What value was reported, on which numerical scale and with what rounding precision? measurement
Normalization and bounds
Captures corrections and the bounds justified by the chosen estimator and input domain.
Attainable range
Establishes how sample size, weighting and normalization affect the coefficient's interpretation.
- Is a finite-sample correction or other rescaling applied, and how is it defined for the weighting scheme used? definition
- What minimum and maximum are attainable under this formula, population size, weight structure and input domain? measurement
Estimation and data treatment Records how evidence becomes a coefficient and how problematic observations are handled.
Grouping, weights and treatment of zero, negative or extreme values can materially change a Gini estimate.
Evidence and calculation
Identifies the input evidence and the calculation path used to produce the estimate.
Input resolution and weights
Distinguishes calculations from unit records, grouped shares and modeled distributions.
- Which dataset, release or publication supplies the estimate, and is it calculated from unit records, grouped distribution shares or a fitted distribution? provenance
- Which weights, interpolation assumptions and calculation implementation are used, and are they documented sufficiently to reproduce the value? measurement
Exceptional inputs
Captures departures from the conventional nonnegative, positive-mean setting and modifications to the distribution.
Zeros, negatives and tail treatment
Records input conditions and transformations that affect mathematical validity or the measured target.
- Are zero or negative values present, what is the sign of the mean, and how does the estimator define or decline to define a result in those conditions? boundary
- Were observations dropped, imputed, top-coded, winsorised or shifted, and what justification and effect on the coefficient are documented? provenance
Uncertainty and comparison Assesses the strength of an estimate and the legitimacy of comparisons made with it.
An apparent Gini difference may reflect estimation error or incompatible measurement choices.
Estimate reliability
Separates statistical uncertainty from sensitivity to coverage and processing assumptions.
Sampling and method sensitivity
Records uncertainty information without treating a sampling interval as a complete account of error.
- What standard error or confidence interval is available, and does its method account for the sample design? measurement
- What evidence shows sensitivity to missing high-value observations, nonresponse, imputation or grouped-data assumptions? measurement
Comparison eligibility
Establishes whether two estimates can support a specified temporal or cross-population comparison.
Harmonization and series breaks
Identifies consequential differences in definitions, coverage and estimation methods.
- Do the estimates use compatible resource definitions, allocation rules, population coverage, reference periods, scales and estimator conventions? boundary
- Should the proposed comparison proceed, require harmonization or remain unresolved given series breaks and uncertainty in the difference? action
Interpretation and permitted inference Defines what the coefficient supports and what requires additional distributional evidence.
A single inequality summary loses information needed for many rankings, explanations and decisions.
Distributional information limits
Relates the coefficient to Lorenz curves and records the limits of scalar rankings.
Lorenz evidence and ranking
Distinguishes a numerical Gini ordering from stronger statements about distributions.
- Are the associated Lorenz curves or distribution shares available, and do the curves cross over the population range? measurement
- Does the proposed claim require only a Gini ranking, or does it require evidence of Lorenz dominance or changes in specific parts of the distribution? boundary
Decision use
Connects the estimate to a stated analytical action without extending its meaning beyond the evidence.
Claim support and complementary evidence
Determines which claims can use the coefficient and which require separate measures or causal analysis.
- What proposed decision or statement will use this estimate, and is there a documented rule for interpreting its value or change? action
- Which additional evidence is required before making claims about poverty, average living standards, subgroup contributions or the causal effect of a policy? action
Evidence and external alignment What the world already says about this thing, gathered so the model can be checked against it.
A model that cannot be lined up against existing standards, identifiers and practice cannot be adopted by anyone who already uses them.
Reported evidence
Findings from the breadth pass, kept separate from the structural claims.
Kinds and varieties
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Market-income Gini (factor or pre-tax, pre-transfer income)
- Gross-income Gini (after transfers, before direct taxes)
- Disposable-income Gini (after taxes and transfers; the usual international headline)
- Wealth or net-worth Gini
- Consumption or expenditure Gini (common where income is poorly observed)
- Household-equivalised versus per-capita or individual Gini
- Generalized Gini / S-Gini (weighted by an inequality-aversion parameter)
- Spatial or concentration Gini (geographic or market shares rather than persons)
- Which of these kinds and varieties hold for the sense of Gini coefficient this model covers, and on what evidence? provenance
Identifiers and schemes
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Wikidata - Q162455 - Item for the Gini coefficient as an inequality statistic.
- World Bank WDI indicator - SI.POV.GINI - Gini index (World Bank estimate), published on a 0-100 scale.
- Eurostat dataset - ilc_di12 - Gini coefficient of equivalised disposable income from EU-SILC.
- OECD IDD series - Gini (disposable income, post taxes and transfers) and Gini (market income, before taxes and transfers) - Named series in the OECD Income Distribution Database, not a compact alphanumeric code.
- Which of these identifiers and schemes hold for the sense of Gini coefficient this model covers, and on what evidence? provenance
Standards and regulation
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Canberra Group Handbook on Household Income Statistics, 2nd ed. (UNECE / Canberra Group) - income concepts underlying most official Ginis.
- OECD Income Distribution Database methodology (OECD) - market/gross/disposable split, household equivalisation, and comparability rules for member countries.
- EU-SILC legal framework and Gini of equivalised disposable income (Eurostat / European Commission) - the EU structural indicator, using the modified OECD equivalence scale.
- Poverty and Inequality Platform (PIP) estimation methods (World Bank) - how survey microdata are turned into the SI.POV.GINI country series.
- Which of these standards and regulation hold for the sense of Gini coefficient this model covers, and on what evidence? provenance
Real-world use
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- National statistical offices and the World Bank publish a country Gini as the headline inequality number in WDI and poverty assessments.
- OECD and EU social reporting compare disposable-income Gini across member states; finance ministries and tax-benefit models (e.g. EUROMOD) report the fall from market to disposable Gini as a fiscal-redistribution statistic.
- Household surveys in many low- and middle-income countries report a consumption Gini because income is seasonal or informal.
- Wealth Gini appears in central-bank and OECD wealth accounts, typically much higher than income Gini.
- Outside income statistics the same formula is used for land-holding concentration, health-outcome or health-expenditure inequality, and market-share concentration.
- Which of these real-world use hold for the sense of Gini coefficient this model covers, and on what evidence? provenance
Typical measurements
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Gini coefficient (unit interval) - 0 (equality) to 1 (maximum inequality); national disposable-income values commonly about 0.24-0.63 - dimensionless (0-1)
- Gini index (World Bank / many NSOs) - about 24-63 for contemporary national disposable or consumption income; global extremes occasionally outside that band - index points on a 0-100 scale
- Wealth or net-worth Gini - commonly about 0.55-0.90 in national household-wealth studies, well above the corresponding income Gini - dimensionless (0-1)
- Redistributive gap (market Gini minus disposable Gini) - often about 0.10-0.25 in high-tax OECD welfare states; much smaller where taxes and cash transfers are limited - Gini points (0-1 scale)
- Which of these typical measurements hold for the sense of Gini coefficient this model covers, and on what evidence? provenance
Failure modes and hazards
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Same Gini, different Lorenz curves: the scalar does not identify the shape of inequality (middle versus tails).
- Survey under-coverage and top-coding of high incomes bias the estimate downward; register or tax-data adjustments can move the number substantially.
- Negative incomes or wealth can push a computed Gini outside [0, 1] or make the mean-normalised index ill-behaved.
- Not additively decomposable by subgroup (unlike Theil / GE): between-group plus within-group Ginis leave a residual, so 'inequality due to X' claims are easy to overstate.
- Cross-country ranking collapses if income concept, equivalence scale, household versus individual unit, or consumption versus income is mixed.
- Finite-sample Gini is downward-biased in small n; naive sample formulae are not interchangeable with population formulae.
- Relative to Atkinson indices with high inequality aversion, Gini is less sensitive to the extreme poor and rich and can miss tail-driven welfare changes.
- Political use as a complete welfare ranking: Gini is not a social-welfare function and does not by itself say whether a rise is 'worse'.
- Which of these failure modes and hazards hold for the sense of Gini coefficient this model covers, and on what evidence? provenance
Regional variation
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Reporting scale: World Bank and many NSOs use 0-100 ('Gini index'); academic and some European series use 0-1.
- EU practice is equivalised disposable income on the modified OECD scale from EU-SILC; the United States long published money-income Gini from the CPS, while CBO series that include taxes and transfers sit closer to the OECD concept.
- Latin America is often described with relatively high income Ginis; Nordic countries with register-based incomes sit at the low end of the OECD range.
- Many African and South Asian official series are consumption Ginis, not income Ginis, and are not interchangeable with OECD disposable-income figures.
- China's national Gini has historically depended on whether urban and rural surveys are pooled and how migrant households are sampled.
- Which of these regional variation hold for the sense of Gini coefficient this model covers, and on what evidence? provenance
Neighbouring kinds and how to tell them apart
Reported by the breadth pass; each item needs checking against its source before it becomes normative.
- Lorenz curve - The Lorenz curve is the full rank-share function; Gini is one scalar functional of it (twice the area to the diagonal). Identical Ginis can hide crossing Lorenz curves.
- Atkinson index - Atkinson is a social-welfare index with an explicit inequality-aversion parameter ε and is more tail-sensitive at high ε; Gini has no ε and a fixed rank-weighted form.
- Theil index / generalised entropy GE(1) - Theil is additively decomposable into between- and within-group terms with no residual; Gini is not.
- Palma ratio - Palma is the income share of the top 10% divided by that of the bottom 40%; it ignores the middle 50% by construction, whereas Gini uses the whole distribution.
- Hoover / Pietra / Robin Hood index - That index is the maximum vertical gap between the Lorenz curve and the diagonal (the share that would need to be transferred to equalise); Gini integrates the gap over the whole curve.
- Percentile ratios (P90/P10, P90/P50) - Ratios use two order statistics only; Gini uses all pairwise gaps. A change confined between two percentiles can move Gini and leave a given ratio unchanged, and conversely.
- Coefficient of variation - CV is the standard deviation over the mean (second-moment / GE(2) related); Gini is based on the L1 mean difference. CV is more sensitive to the upper tail.
- Gini impurity (CART / classification) - Gini impurity is 1 − Σ p_k², a node-heterogeneity score in decision trees. It shares a quadratic form with some Gini algebra but is not an inequality index of a cardinal distribution.
- Which of these neighbouring kinds and how to tell them apart hold for the sense of Gini coefficient this model covers, and on what evidence? provenance
Sources
- Variabilità e mutabilità - Original definition of the mean-difference / concentration measure later called the Gini coefficient.
- On the measurement of inequality - Places Gini among Lorenz-consistent inequality indices and contrasts it with the Atkinson family.
- Canberra Group Handbook on Household Income Statistics, 2nd edition - International conventions on income concepts (market, gross, disposable) that determine which Gini is being reported.
- World Development Indicators: Gini index (World Bank estimate), SI.POV.GINI - The dominant country-year series, the 0-100 reporting scale, and typical cross-country ranges for income Gini.
- OECD Income Distribution Database (IDD) - The market versus disposable split, equivalisation practice, and OECD-area levels used in policy comparison.
- Measuring Inequality, 3rd edition - Formal properties (Pigou-Dalton, decomposability limits, relation to the Lorenz curve) and neighbouring indices.
What the second pass must settle
- Does the registry intend this entry to cover distributional Gini coefficients across all quantities, or specifically socioeconomic inequality?
- Does an existing Vercy world model already own the distributional Gini concept, requiring this registry entry to link to it?
- Which estimator and finite-sample normalization conventions should the model explicitly recognize, particularly for unequal weights?
- How should estimates involving negative observations or nonpositive means be classified, and which source-backed interpretations should be accepted?
- Which authoritative methodological sources should establish comparison requirements for income, wealth and consumption estimates?