WHY A UNIVERSITY LAB

Your team built the intelligence. A university lab tests what makes people choose it.

A company is built to align people and ship. A university lab can convene people who would rarely fit one reporting line — poets, fashion scholars, designers, behavioral scientists, clinicians, statisticians, and software engineers — and give their disagreements a scientific process. The point is not eclectic brainstorming. It is independent evidence for the customer and business decision you need to make next.

Independence that protects your decision

Your team knows the product, market, and constraints. The lab is useful because it can test the attractive idea without needing it to win. A null or negative result can save more value than a flattering story.

A process designed to lose an argument

We state the null hypothesis, primary outcome, meaningful effect, sampling plan, exclusions, and decision rule before the result. We estimate uncertainty, correct multiple tests, try to falsify the claim, and replicate before scale.

A team no corporate org chart should have to absorb

The unusual team is deliberately temporary and problem-specific. Your product organization keeps moving; the lab brings in the exact intellectual friction a consequential decision needs, then turns it into evidence your organization can use.

A lab earns its place by being willing to falsify the attractive idea.

We do not sell predetermined conclusions or promise growth. We decide in advance what result would support the intervention, what would count as too small to matter, and what would tell us to stop. Positive, null, and negative results can all strengthen a real business decision.

Their first loyalty is to people, knowledge, or craft — not the quarter.

That is a feature. A poet may care about what an answer does to dignity; a clinician about safe action; a designer about who is excluded; an engineer about reliability and cost. Revenue, adoption, and retention still matter because your business must decide — but they are measured alongside the human consequence, not instead of it.

Productive incompatibility is part of the research instrument.

These people may disagree about what the problem even is. We do not force an artificial consensus. We convert competing readings into rival hypotheses, and the research contract — not seniority or presentation skill — decides what survives.

Poet or literary scholar

Sees voice, sequence, metaphor, omission, and the implied reader — including what the answer quietly asks a person to believe or do.

Fashion and identity scholar

Sees status, belonging, taste, cultural signals, and who feels addressed or excluded before a conventional usability metric notices.

Designer and HCI researcher

Sees hierarchy, affordance, cognitive load, modality, and the moment an intelligent answer becomes difficult to use.

Behavioral scientist or clinician

Sees confounds, stress, motivation, appropriate action, harm, and the gap between what people say and what they actually do.

Statistician or methodologist

Asks whether the effect exceeds noise, whether the analysis was chosen after seeing the answer, and whether one lucky comparison is being mistaken for a finding.

Software or LLM engineer

Makes the intervention reproducible, isolates model and interface changes, and keeps latency, reliability, tokens, and serving cost inside the decision.

Humanities do not decorate the answer. They generate variables worth testing.

Close reading, narrative theory, cultural analysis, fashion studies, and design research reveal mechanisms a benchmark may miss. We operationalize the relevant insight as a controlled intervention and measure what real people understand, choose, pay for, use, remember, or safely act on. Not every human experience should be reduced to a number; any product claim made from it should survive an empirical test.

Literature and poetry

WHAT IT SEES
Voice, order, metaphor, narrative distance, and omission can change what an identical set of facts means to a reader.
HOW WE TEST IT
Render matched answers with competing narrative structures while holding model, facts, and task constant.
WHAT WE MEASURE
Comprehension, recall, calibrated belief, appropriate action, message preference.

Fashion, identity, and culture

WHAT IT SEES
Visual and linguistic signals communicate status, belonging, expertise, intimacy, and exclusion before a customer evaluates capability.
HOW WE TEST IT
Vary presentation, persona, visual language, and cultural cues across customer segments without changing the underlying answer.
WHAT WE MEASURE
Perceived relevance, trust calibration, choice, willingness to pay, subgroup effects.

Design and human–computer interaction

WHAT IT SEES
Information hierarchy, progressive disclosure, modality, and control determine whether capability is usable at the moment of decision.
HOW WE TEST IT
Compare interaction structures with the same content and model under controlled tasks and realistic time pressure.
WHAT WE MEASURE
Task success, time, error, abandonment, confidence calibration, support demand.

Psychology and behavioral science

WHAT IT SEES
Attention, emotion, norms, incentives, and cognitive load can make a technically correct answer ineffective or unsafe.
HOW WE TEST IT
Predefine a behavioral mechanism, comparator, and manipulation check, then test it with the people and context that matter.
WHAT WE MEASURE
Observed behavior, adoption, retention, safe escalation, decision quality, heterogeneous effects.

From a strange insight to a decision you can defend.

The method is familiar science, applied to human–LLM interaction and business decisions. A p-value is one diagnostic — never a truth machine. Effect size, uncertainty, design quality, practical importance, and replication decide whether the result deserves action.

  1. 1. Observe and interpret

    Interviews, close reading, field observation, cultural analysis, and product data locate the human mechanism worth studying.

    UCL — Methods for HCI Research
  2. 2. State what could be false

    We define the null hypothesis, rival explanation, primary outcome, and minimum effect that would be large enough to change the business decision.

    Center for Open Science — Registered Reports
  3. 3. Commit before seeing the answer

    Population, sample, exclusions, stopping rule, intervention, comparator, outcomes, and confirmatory analysis are fixed before results can tempt us to move the goalposts.

    Center for Open Science — Preregistration
  4. 4. Test the real system with real people

    We control the model, prompt, interface, and exposure where possible; randomize or match conditions; document the deployment context; and include ethics and risk in the design.

    NIST — AI Risk Management Framework
  5. 5. Estimate — do not merely declare significance

    We report effect sizes, intervals, p-values where appropriate, practical thresholds, and uncertainty. A small p-value does not prove that an effect is true or important.

    American Statistical Association — Statement on p-values
  6. 6. Correct the temptation of many comparisons

    When several confirmatory hypotheses are tested, procedures such as Holm correction control the family-wise error rate instead of rewarding the luckiest result.

    Holm (1979) — A Simple Sequentially Rejective Multiple Test Procedure
  7. 7. Try to make it fail again

    A promising effect is replicated across a new sample, language, model, workflow, or setting before it becomes a general claim or a large deployment.

    Many Labs 2 — Replicability across samples and settings

Your team should not become a university — and the lab should not pretend to run your business.

Trying to recreate this mix permanently inside one reporting structure either neutralizes the dissent or burdens the product organization with people it was never designed to manage. A finite university partnership preserves both strengths: your team keeps its speed and judgment; the lab supplies independent friction, method, and evidence.

YOUR TEAM

Your company brings

  • Product, customer, and market knowledge
  • The real decision and its economic stakes
  • Access to workflows, users, and operational constraints
  • Risk owners, domain experts, and implementation capacity
  • The judgment and authority to act on the result
THE LAB

The university lab brings

  • Independence from the preferred internal answer
  • Humanities, design, science, medicine, and engineering in one study
  • Human-participant ethics and transparent procedures
  • Null hypotheses, controls, thresholds, uncertainty, and stop rules
  • Falsification, multiple-testing correction, and replication

This is not a new scientific method. It is an established university pattern — aimed at LLM customer value.

Universities have long built centres around problems that no single discipline can own. These are precedents, not affiliations: each demonstrates a part of the method we adapt for human-centric LLM research with measurable customer and business outcomes.

Stanford University

Stanford Institute for Human-Centered AI

An interdisciplinary AI institute spanning Stanford’s seven schools, guided by human impact and the aim of augmenting people.

Visit the centre
Massachusetts Institute of Technology

MIT Media Lab

An interdisciplinary lab built around unconventional mixing of seemingly disparate areas — including AI, art, design, health, robotics, music, and human–machine interaction.

Visit the centre
Arizona State University

Center for Science and the Imagination

Brings writers, artists, and other creative thinkers together with scientists, engineers, and technologists, connecting human narratives to scientific questions.

Visit the centre
University College London

UCL Interaction Centre

Studies people, technology, and society through the scientific traditions of computer science and human sciences, at the intersection of engineering, behavioral science, and design.

Visit the centre
Stanford University

Stanford Literary Lab

A research collective that applies computational criticism to literature — a direct precedent for humanistic questions investigated with reproducible analytical methods.

Visit the centre
Princeton University

Center for Digital Humanities

Treats computation, data science, and the humanities as reciprocal partners, combining critical interpretation with data-intensive research and software engineering.

Visit the centre
University of Oxford

Institute for Ethics in AI

Brings philosophy, humanities, and STEM researchers together with leaders in technology, business, government, and civil society to provide independent perspectives on AI.

Visit the centre
University of Maryland

Human–Computer Interaction Lab

A long-running interdisciplinary HCI lab connecting information studies, computer science, psychology, education, English, engineering, journalism, and other fields.

Visit the centre
Harvard University

metaLAB

A knowledge-design lab rooted in arts and humanities, working across code and narrative, data and ethics, networks and archives.

Visit the centre
London College of Fashion, UAL

Fashion Innovation Agency

Connects fashion expertise, brands, and technology partners through proofs of concept and prototypes, then returns that learning to the university.

Visit the centre
Royal College of Art

Helen Hamlyn Centre for Design

A people-centred inclusive-design research centre working with business, government, academia, and the third sector across design, psychology, neuroscience, health, and systems thinking.

Visit the centre
RESEARCH CONTRACT

Before we test, we agree what would change the decision.

This is the compact research contract behind every sponsored study. It protects your team from an impressive result that answers the wrong question — and protects the lab from quietly redefining success after the data arrive.

01Business decision, owner, and customer behavior at stake
02Humanistic or behavioral mechanism and rival explanation
03Falsifiable hypothesis and explicit null hypothesis
04Population, sample-size rationale, recruitment, and exclusions
05Intervention, comparator, model version, and deployment context
06Primary outcome, minimum meaningful effect, and secondary outcomes
07Analysis plan, uncertainty reporting, and correction for multiple tests
08Rules for positive, null, negative, and ambiguous results
09Replication condition required before scale

Bring the product judgment. We will bring the people your org chart cannot — and a method strong enough to test them all.

Start with one consequential customer or business decision. We will assemble the smallest genuinely heterogeneous team that can challenge it, then design a study whose result remains useful even when the original idea does not survive.

Start a rigorous project

What customer or commercial decision should the evidence strengthen?

You know the opportunity. Tell us the decision and choose the outcomes that would make the result useful.

Your role
What should the study help you measure?
Draft the conversation

Your email app opens with the draft; this site stores nothing.