All experiments

Experiment 02 / Logosmose / 2026

Can AI-assisted evaluation earn trust?

A written-reasoning experiment built to make the question, constraints, rubric and result inspectable rather than asking users to trust a single opaque judge.

FounderProduct designSolution architectureLead full-stack development

Evaluation can be structured without pretending it is objective.

Personal experience with inconsistent human evaluation made the trust problem concrete. The defensible product claim became narrower: evaluate a written response against a public rubric, not judge a person’s intelligence.

Make the test brief, constrained and inspectable.

  1. 01

    Reveal one difficult question.

  2. 02

    Write for five minutes, within 2,000 characters.

  3. 03

    Evaluate the response across five public criteria.

  4. 04

    Return the score, criterion breakdown and rank.

Probabilistic help. Explicit accountability.

AI leverage

Where models contribute

Four model providers contribute to grading. Their outputs are combined around a versioned rubric covering relevance, reasoning, clarity, accuracy and concision; no single model controls the verdict.

Human judgment

Where responsibility stays

I defined the product claim, rubric, constraints, aggregation rules and acceptable language. Users decide whether the result is useful; the system does not claim to measure intelligence.

The constraint and the verdict are visible.

These production screens show the core loop now available for free participation and training. They demonstrate interface behavior, not user traction or grading validity.

Logosmose answer screen showing a five-minute timer, a revealed question and a 2,000-character response limit.
Production UI / timed response with explicit limits
Logosmose result screen showing an overall score, five-criterion chart and rank.
Production UI / score, criteria and rank returned together

The original financial dimension created legal exposure, gambling associations and payment-processor friction. I removed it. That decision left a smaller free loop capable of testing trust and repeat participation before adding commercial complexity.

Next test

Whether the free core loop creates voluntary repeat use.