Blindstairs
Choose language. Current language: English
Merified HRISRequest demo
A human stone figure held intact while transparent geometric layers reveal its underlying structure

Our position

AI should expandopportunities

Explore our edge

About Blindstairs

Equity is not a feature. It is the reason we build.

We believe every person, no matter their background, deserves a fair shot at entering, moving, and growing at work. Technology should help people progress while accountable human judgment shapes their future.

BlindStairs is an R&D-led company. We turn research questions about bias, repeatability, and accountability into deployable product architecture that organisations can inspect and govern.

The problem we chose to work on

AI can scale a decision. It can also scale what should never influence it.

Across People Operations, consequential decisions shape who gets to enter, progress, and thrive at work. They also carry the biases of the environments in which they are made.

General-purpose AI learns from historical data and the wider world. When it is asked to evaluate a person directly, gender, age, nationality, ethnicity, or indirect proxies can influence its output even when they are irrelevant to the decision.

If the professional evidence stays the same but a changed name can change the score, automation has not removed the problem. It has made it faster and harder to inspect.

01Same professional evidence
02One irrelevant signal changes
03A different result

An irrelevant signal should never change professional judgment.

Controlled name-swap experiments

Same professional evidence. Different names. The score should not move.

Within each model test, we held a base CV constant, changed only the name set, and repeated the evaluation. Direct LLM scoring produced dispersed results and a shift between name groups. BlindStairs separated model interpretation from deterministic aggregation.

Representative runs shown · broader invariance validation across 10,000+ paired CV comparisons
Fixed within each testOne base CV
Changed variableName set only
Repeated runs50 per group
Read correctlyCompare within each panel

Direct LLM screening

Variable in this test
General-purpose model

Gemini 2.5 Pro

1.78 mean shift
Gemini 2.5 Pro score-density distribution by name groupA separate controlled run with fifty evaluations per name group. Within this model test, the fixed professional evidence produced dispersed distributions and a mean shift of 1.78 points when the name set changed.
Male-coded mean 51.53Female-coded mean 53.31
General-purpose model

GPT-OSS-120B

1.72 mean shift
GPT-OSS-120B score-density distribution by name groupA separate controlled run with fifty evaluations per name group. Within this model test, the fixed professional evidence produced dispersed distributions and a mean shift of 1.72 points when the name set changed.
Male-coded mean 59.51Female-coded mean 61.23

The same professional evidence produced a range of scores. Only the name changed.

BlindStairs

Stable by design
Male-coded nameFemale-coded name
BlindStairs repeated-score resultThe BlindStairs pipeline returned 64 out of 100 for every repeated run in both name groups.
64/100in every repeated run shown

Irrelevant signals are removed before evaluation. Requirement-level results are aggregated deterministically outside the LLM.

Method. Each external-model panel is a separate controlled test. Within that panel, a fixed base CV was evaluated 50 times per name group, with only gender-coded names changed. Absolute scores should not be compared between model panels. The broader validation programme covers more than 10,000 paired CV comparisons.

Interpretation. The experiment shows repeatability and sensitivity to a changed signal under the stated conditions. It does not, by itself, measure real-world discrimination. BlindStairs is engineered so irrelevant sociodemographic signals do not determine the calculated result, while final judgement remains human.

What the research taught us

Bias is not a setting you can switch off.

Better data, anonymisation, and auditing all matter. None is a credible product guarantee that a general-purpose language model will become neutral or deterministic.
01

Diversify the training data

Better representation matters, but historical patterns, structural inequality, and indirect proxies can remain embedded in the model.

02

Hide protected attributes

Removing direct identifiers helps, but other details can still act as proxies. Strip too much and valuable professional context disappears too.

03

Audit the outputs

Audits can reveal failures. They do not make direct LLM scoring deterministic, nor do they stop the next output from changing.

We stopped asking a general-purpose model to be unbiased. We redesigned the decision so its output could not control the score.

Our architecture

Understanding inside the model. Decision logic outside it.

Language models help interpret complex evidence. Explicit standards, deterministic aggregation, and accountable people govern the result.
01Remove irrelevant signals

Identity and non-pertinent sociodemographic cues are stripped before evaluation.

02Make the standard explicit

The decision is broken into simple, weighted criteria agreed by people.

03Use LLMs to read evidence

Models interpret source material criterion by criterion, never as a final scorer.

04Calculate outside the model

Deterministic rules aggregate the result for accountable human review.

Human authority is not another stage. It governs the whole process.

Research · partnerships · applied innovation

Human-centred AI needs more than good intentions.

It needs evidence, accountable architecture and collaboration across research, policy and practice. If that is your work too, we would like to hear from you.

Talk with us

Backed by

  • Women TechEU
  • EPIC-X, funded by the European Union
  • AI on Demand
  • CDTI Innovación
  • Universitat de Barcelona
  • ACCIÓ, Generalitat de Catalunya