Here is a hard problem that looks easy. Ask thousands of people what freedom and equality mean to them, then find what they agree on. Polls can measure agreement, but only with statements someone wrote in advance. Focus groups draw out what people actually think, but only for a dozen people at a time. Neither tells you what a country agrees on in its own words.

Jigsaw, the unit at Google that works on technology and society, published a research note on an attempt to do both at once. It's worth reading closely, because it's a clear example of what AI is good for in collective deliberation, and of where the hard questions move once it works.

What they did

In the We the People pilot, run with the Napolitan Institute, more than 2,400 Americans drawn from all 435 congressional districts answered open-ended questions about freedom and equality, and responded to each other's ideas. Together they produced more than 1.6 million words.

Jigsaw's sensemaking tools, built on Gemini and released as open source, then did three jobs: sort the responses into topics and opinions, summarize them, and generate new short statements likely to win broad agreement given what participants had already said. People chose the final set. The result was 26 plain statements, put back to participants to agree or disagree with. Most drew more than 80% agreement, and 94% of participants said they felt their opinion was represented.

That's a remarkable outcome for a topic most people would expect to be divisive. The research question was whether AI could compress the breadth of a national conversation into a few statements people actually endorse, and the answer was yes.

The part that is new

The novel step isn't summarization. It's proposition generation: the model writing statements nobody said, chosen for being likely to be endorsed by the people who said everything else. That puts the model in the position of a very good mediator, someone who listens to a divided room and says “it sounds like all of you believe this,” and is right.

Two design choices kept that role honest. Participants didn't just receive the statements; they voted on them, so every claim of consensus was tested against the people it described. And a person made the final selection instead of letting the model's ranking stand as the answer. The model proposes; people decide.

Where the hard questions move

Once the machinery works, the difficult questions stop being technical and become questions about meaning. This site's vocabulary helps here.

Agreement is the ; meaning is the . A statement with 85% agreement has a large, measurable extension: the set of people who ticked “agree.” What it doesn't tell you is whether they agreed with the same thing. “Everyone should have an equal chance to succeed” can be endorsed, sincerely, by people who would design opposite policies to deliver it. Broad agreement on a sentence isn't broad agreement on what it means.

Optimizing for agreement has a direction. If the objective is “statements likely to be endorsed,” the easiest way to raise the score is to make the statements more general. Abstraction buys agreement. That's 's law in a new setting: the measure (agreement) and the goal (finding what people genuinely share) coincide until you push on the measure. Human selection and participant voting are the right counterweights, and it's worth being explicit that they are the part doing that work.

The mediator holds the pen. Whoever writes the candidate statements shapes which agreements are even available to find. That isn't an accusation; every facilitator has this power. But it's the same concern we have with multi-agent systems, where the that frames a question steers the answer. The structural answers carry over: keep the framing step visible, let the people described test every claim about them, and keep a person accountable for the final choice.

What this suggests for building with AI

The pilot is about civic deliberation, but its shape is general, and it's close to how we try to use models in our own work:

Finding common ground, carefully

It's easy to be cynical about AI and democracy, and this project is a good counterexample: a careful use of models to help a large group hear itself, with people kept in the loop at the points that matter. The open question it leaves isn't whether AI can find agreement. It can. It's whether the agreement it finds is about meaning or only about wording. The measurements can't answer that alone. Deciding it will always involve people reading the statements and asking what they are for.