Argupedia

The Policy Encyclopedia

Our Method

An argument database, built for the way we think.

The following text is our reasoning for why we built Argupedia. This section, too, is still in beta. We are happy to answer questions — there are far more ideas and thoughts in this concept than fit on this page.

The Problem

Deciding is messy.

Democracy rests on a promise: arguments compete openly, and the best one wins. That is the precondition for a democracy making the right decisions and implementing the best policies. This promise is not reliably kept. Our brain responds to the ethos and pathos of a speech, to identity and social signals — to things that say little about whether a policy actually works.

On top of that, we often use only the information that is easily available and stop looking. From it we build a coherent story on which our position then rests. Daniel Kahneman gave this pattern a name: What you see is all there is (Thinking, Fast and Slow, 2011).

As long as information was scarce, that was a reasonable bet: you work with what you have. Today information is no longer scarce. But more information does not automatically produce better decisions. Our working memory holds only about four units at a time (Cowan, 2001) and is quickly overloaded, and we preferentially reach for whatever supports the opinion we have already formed (Nickerson, 1998). So decisions — democratic ones in particular — continually fall short of what would theoretically be possible. And in the end, a democracy is only as good as the decisions it produces.

Why hasn't this been solved long ago?

If the cognitive sciences understand these phenomena so precisely, why is the problem still unsolved? For one thing, it does not feel like a problem. Our brain reports back a good, considered decision even when it hinged on simple signals. That we do not decide as well as we believe usually only shows in the outcome.

For another, it is not true that the problem is unsolved. It is only unsolved in democracies, so far. Aviation solved it — with technology, procedures and methods that put the human being at the centre. At the start of the jet age, the rate of fatal accidents in commercial aviation was above 35 per million flights; today it is around 0.1 (Boeing Statistical Summary). At today's volume of roughly 40 million commercial flights a year, the old rate would mean well over a thousand fatal crashes — every year. Averaged over the past five years, there were six worldwide (IATA).

Where it comes from

How we think.

We decide on signals

A cue is anything that reaches us in a decision and enters the judgment. Not just the spoken argument, but also the speaker's face, the party behind the proposal, the tone of the coverage, the feeling a story triggers. The term comes from the psychologist Egon Brunswik (1956) and is deliberately broader than the everyday word signal, because it also includes what we do not even perceive as information.

What cues would have to be measured against

Almost every political decision is a choice between two futures: F0 without the measure, F1 with it. What should carry the decision is the difference between the two.

better ↑ worse ↓ Time today F0 F1 Δ
F0: the future without the measure F1: the future with it Δ: what it actually changes

Cues would have to be measured against this difference. Some do that well, others hardly at all. How well a cue indicates which future is better is what we call its validity. Whether we check its validity at all is a second question, and how heavily it then counts in the judgment a third.

Validity

How well does the cue carry the answer? A robust study on the measure's effect is highly valid. The speaker's demeanour says almost nothing about the difference between the futures.

Processing

Do we check it? The fast, intuitive judgment takes the cue as it feels. The slow, examining one asks whether it holds. That is effortful and rarely kicks in on its own.

Usage

How heavily does it weigh in the judgment? The weight the cue ends up getting — regardless of whether it deserves it.

Validity and usage should coincide. What carries the answer well should weigh heavily; what does not carry it, lightly. In practice they routinely come apart (Hammond et al., Social Judgment Theory), and that hangs above all on the middle column: unchecked cues often weigh as much in the judgment as checked ones, sometimes more, because they are catchier. Checking happens too rarely. And from the inside, both feel the same — like careful thinking (Wilson & Brekke, 1994).

The Lens Model · Brunswik & Hammond

Validity is not weight

Every cue has two strengths: how well it actually reflects reality — and how heavily the judgment leans on it. Political judgment fails where the two come apart.

✓ Validity REALITY the right decision DIRECT CUES INDIRECT CUES Evidence Logical structure Documented effect Single study · contested Speaker confidence Party affiliation Emotional charge Group consensus JUDGMENT the decision HIT RATE 38 % Correction · Calibration
thin = low validity thick = heavy usage (weight) direct cues indirect cues

The consequence over time. Where judgment hits poorly, the trajectory stays far below what would be possible. GDES shifts it measurably towards its potential — not least because errors become visible through the correction loop.

Value Time Potential intuitive GDES

The primed receiver

Cues do not land on an unbiased person, but on one with a history. Whoever brings the necessary knowledge can place a complex signal. Whoever lacks it is overwhelmed by the same signal, or it barely helps them.

More often still, people arrive with an opinion already formed and argue onward in its direction instead of testing it (Lodge & Taber, 2013). And there are cases in which group and identity signals decide more about how a judgment turns out than the facts that belong to the matter (Kahan, 2013). One and the same number is then believed or torn apart depending on who cites it.

We all know this

There are by now countless studies describing human reasoning errors, shortcuts and blind spots. You do not need them to recognise the phenomenon. Perhaps you have defended a conviction you never really examined. Perhaps you have passed on a number without knowing where it came from. Perhaps the opinion came first and the arguments afterwards. All of this has happened to us too — and building this site makes it especially visible.

Conditions that fit human beings

Evaluating a political measure and its effects over years is an extraordinarily complex task. Many cues can hardly be placed without expertise. Much runs in parallel. You have to decide which argument even belongs to the matter, how certain it is and how heavily it should weigh. That is more trade-offs than a person can hold in their head at once, and under overload the quality of judgment does not drop a little — it drops sharply (Sweller, Cognitive Load Theory). That is exactly where we fail without the right tools.

The individual error cannot be predicted. We only know that it happens, what kind of errors to expect and where they arise. From that follows what a tool must deliver:

  • Make the right cues visible. Show which cues speak to the matter — to the difference between the two futures.
  • Lay validity and weight open. For every cue, make traceable how well it carries and how heavily it counts, so both become comparable.
  • Present without overload. Summary at a glance, details on demand — otherwise the sheer volume of information crushes exactly the examining thinking it is meant to equip.

And each individual decision is rarely just good or bad. Usually there are several paths — a clearly better one, a slightly better one, one that is barely distinguishable. And they add up: the distance to what would be possible is the sum of the non-optimal decisions. The next section builds a procedure out of this.

The Solution

A standard for debates.

We focus the debate on what it decides, and we present exactly that. The choice is between F1, the world with the measure, and F0, the realistic course without it. At its core, every argument claims a difference between these two futures — a delta. We score it in three quantities: Impact, Plausibility and Value. We have worked this procedure out as an open standard, the Global Debate Evaluation Standard — GDES for short.

  1. Impact: naming the delta

    Every argument claims a difference between F1 and F0: fewer emissions, higher rents, shorter waiting times. The first step makes this claim explicit and measures its size — the Impact.

  2. Plausibility: how certain is the delta?

    Predictions are uncertain. A large delta with weak evidence must not count like a smaller one with strong evidence. So the impact is weighted by its probability of materialising — its Plausibility. Arguments about credibility, feasibility or side effects enter here; they do not disappear, they change the delta's impact and plausibility.

  3. Value: how much is it worth to us?

    Not every dimension matters equally to us. The expected effect is weighted by its Value — by how much what is affected matters to us. This is where values enter, openly and visibly, instead of hiding inside seemingly factual claims.

Summed up in one formula:

Argument Score
Score = V × I × P ÷ 10
V: Value  ·  I: Impact  ·  P: Plausibility

The formula is only part of the tool. Its real use lies in being able to zoom in and out. Every single rating — every V, every I, every P of every argument — remains separately visible, contestable and correctable. Added together, the scores of all pro and con arguments yield a measure's balance: a fast overview of a complicated debate, behind which the full reasoning can be unfolded at any time. Exactly the form that examining thinking needs: summary at a glance, details on demand.

Why a platform

How we implement it here.

You can apply the standard to any single debate. A platform can do more. It keeps Value and Impact on the same scale across all debates and learns from it. You can see how the ratings relate to each other, how the scale shifts over time, and how a change would play out across the results of entire debates.

One shared scale

Deltas come in different units: tonnes of CO₂, euros, life years, waiting time. To make them comparable, they have to be normalised. We do this in welfare euros per year, and we derive every number from sources and disclosed calculation. The scale is open at the top — a large effect may look large, a small one stays small. So that the numbers in a single debate remain readable, we calibrate the impact to the topic at hand. This factor only shifts the yardstick, like switching from metres to kilometres. The ratios between the arguments stay the same, and both views can be converted into each other at any time.

We try to keep this normalisation constant across all debates. When we notice that a rating is off, we can adjust it. We are still in the learning phase on this question, and we document openly where things stand.

Values and evidence

A global value register. Eight classes, from 10 down to 3. Life and health stand at 10, the constitutional core at 9, basic security and educational opportunity at 8, participation and the environment at 7, functioning systems and prosperity at 6, money at 5, comfort at 4, marginal benefits at 3. The 5 is the reference point of the whole scale: every cost argument hangs on it, and a euro spent weighs the same in every debate. The register is identical for all debates — no debate can talk up its own importance.

Evidence with ceilings. No number stands without a source, and every source carries its type visibly on the argument: scientific studies, direct precedents, clear mechanisms. Weaker foundations, such as projections without disclosed methodology, are allowed but capped. They cannot claim high certainty.

All assumptions in the open. The normative settings — say, the value of a healthy life year or the weighting of income — are publicly documented in the Registry; the derivation of the procedure is in the Whitepaper. They are only ever changed centrally, never case by case and never with a running debate in view.

AI — with humans in the loop

Artificial intelligence produces fluent, persuasive text at a scale no individual can check. We use it ourselves: it drafts the initial evaluations. Because of the way we use it, the result is not a persuasive text but a transparently evaluated debate in which every single point can be verified. Every number must be derived — with source, calculation and assumptions — and every derivation can be attacked at exactly one place. Not "the evaluation is wrong", but "the impact of argument 3 is set too high, here is the better source". That is how judgment stays with the human.

We ourselves check the analyses primarily for obvious errors and fill in the coarse gaps we notice. Beyond that, we count on the community, on better AI systems and on what we learn from our own work. In that combination we trust ourselves to improve the evaluations step by step.

Beta phase

Argupedia is in beta. We are learning very fast right now, and larger changes can happen at any time. Take the evaluations as a starting point, not as the last word. And that is exactly why we need your feedback: we are glad about every message to [email protected].

Our goal

We want to present debates in a form that makes it easier for us humans to understand them, to check them and to draw our own conclusions. And we hope that the focus on evaluating real change, together with this kind of transparency, also influences how debates are conducted. We do not consider that a nice extra, but a democratic imperative.

Why more is at stake

The democratic imperative.

A democracy draws its legitimacy from two sources. One is the process: the people take part in deciding. The other is the outcome: the decisions are good for the people who make and carry them. Both sources depend on how well debates are conducted.

Process: understanding and co-deciding from: casting a vote  →  to: truly grasping the decision Outcome: good decisions low legitimacy high legitimacy
Legitimacy is not a switch but a matter of degree. It grows along both axes at once.

The process goes beyond the right to vote. A vote with no judgment of one's own behind it is hardly an act of self-determination. Whoever takes part in deciding should be able to understand why they decide the way they do — and also why decisions are made about them the way they are, so that they can take this into account at the next election. Kant defined enlightenment as the exit from self-incurred immaturity — as the ability to use one's own understanding. He was not a democrat in our modern sense, and his question concerned the mature use of reason, not the ballot box. But read democratically, it means: a democracy owes its citizens not only the right to vote, but the conditions under which a judgment of one's own can arise in the first place.

These conditions are middling today. Many of the formats in which we negotiate politics are designed to work fast, not to be understood. Political statements are often catchy without delivering the transparency that would let you trace why someone is for or against a measure. And where things are not sharpened to a slogan, the opposite often happens: you are left alone with unordered complexity that overwhelms examining thinking. In both cases, your own judgment falls by the wayside.

Jürgen Habermas formulated a condition for legitimate democratic decisions: that only the unforced force of the better argument prevails — that nothing tips the scales except the matter itself.

That the best argument wins.

The second source of legitimacy is the outcome, and much is open there too. Many problems in Germany, large and small, have been known and unsolved for years. That gnaws at trust in democracy — and trust is not won back through better communication, but through better decisions. Both hang on the same place: on the way we debate. Whoever recognises the better proposal more reliably also chooses it more often.

From this follows the democratic imperative. If people decide the way the research describes, and if legitimacy depends on citizens being able to form a judgment of their own and on good decisions standing at the end, then preparing information so that a human being can understand and evaluate it is not a friendly extra service. It is the precondition for a decision keeping its own promise.

The Athenian assembly opened with a question to everyone present: Who wishes to speak? That was the promise of democracy in its oldest form. Everyone may. Today that is no longer enough. That everyone may speak only has value if those who listen can also judge what is being said. That is exactly what Argupedia tries to make possible.

Threshold Gradation Conditions for independent judgment → Legitimacy →

Not a threshold — a gradation.

The better the conditions, the more legitimate the decision. From this follows the ought: a democracy should provide its citizens with the best information in the most comprehensible form — so that the decisions that emerge from it carry a higher legitimacy themselves.