Can instruction reinforcement keep LLMs aligned under context pressure?
Test whether strategically repeated or structurally surfaced constraints improve compliance in long, multi-turn and tool-using LLM tasks without materially increasing tokens.
View hypothesis