All projects
seeking partner

Can instruction reinforcement keep LLMs aligned under context pressure?

Test whether strategically repeated or structurally surfaced constraints improve compliance in long, multi-turn and tool-using LLM tasks without materially increasing tokens.

HYPOTHESIS

Reinforcing critical instructions at decision points will improve specification compliance by at least 8 percentage points while increasing total tokens by no more than 5%.

PRIMARY METRIC

Rate of tasks satisfying all predefined critical constraints.

MEANINGFUL THRESHOLD

+8 percentage points instruction compliance with no more than 5% additional tokens.

BUSINESS TARGET

More reliable LLM and agent behavior without retraining or a material inference-cost increase.

DECISION RULES

Compliance gain meets both thresholds

Proceed to a larger cross-model replication and product pilot.

Compliance improves but token budget is exceeded

Optimize reinforcement frequency and placement.

No meaningful improvement

Do not deploy repetition as a general reliability mechanism.

Business outcomes are research targets, not guarantees. A null or negative result may still create substantial value by preventing investment in an ineffective product, feature, or campaign.

Population

Representative multi-turn tasks with persistent constraints, including tool-using and agentic workflows.

Intervention

Critical instructions restated or structurally surfaced at selected context boundaries.

Comparator

The same task with each instruction provided only once at its original position.

Secondary metrics

Task accuracy · Total tokens · Latency · Constraint-specific failure rate

Does this hypothesis match a customer or business decision your company must make?

Discuss this project

Discuss a Research Project

Describe where the uncertainty sits and what you would like to test. Your email app will open with a structured draft; the site stores none of this information.

Open email draft