All projects
seeking partner

Do long-running LLM agents still pursue the user's original goal after dozens of steps?

Measure goal drift across long-running LLM agent trajectories and test lightweight checkpoints that restore user intent without restarting the workflow.

ProductAdoption & Retention
LLM Software & Platforms
WHAT YOU CAN EXPECT

Evidence protects you in both directions.

01
Illustrative result · threshold met

−31% · goal-drift failures

Customer success managerIllustrative result · threshold met
−31%

goal-drift failures

Intent checkpoints triggered by detected plan drift kept long-running agents aligned with the customer's original goal more often without restarting the workflow.

What your study would pin down

You get the drift triggers for your workflows, the checkpoints that actually recover customer intent, and the point where intervention costs more than the failures it prevents.

Decision

Trigger checkpoints only when drift signals indicate that customer intent is at risk.

Public precedent · not your resultApollo Research / AIES

Scaffolding kept the best agent nearly perfectly on-goal beyond 100,000 tokens.

AIES 2025 experiments found some goal drift in every evaluated model, while a scaffolded Claude 3.5 Sonnet maintained nearly perfect goal adherence for more than 100,000 tokens in the hardest setting.

Source: AAAI/ACM AIES 2025 ↗The precedent shows the effect can exist elsewhere. Your study determines whether, where, and how strongly it holds for your customers, workflows, and economics.
02
Decision

Evidence protects you in both directions.

Customer success managerThreshold not met · do not scale
What the data may show

No meaningful drift reduction

If the intervention does not clear the predefined threshold, that is evidence against spending more to build, launch, or scale it in this context.

Decision

Do not add checkpoints as a general orchestration mechanism.

Customer success managerThreshold met · value left unused
The other expensive error

A real opportunity can still be left on the table.

If the intervention clears the threshold but the business keeps the current approach, measurable savings, revenue, adoption, or risk reduction may remain unrealized.

Decision

Act only when the measured opportunity is large enough to justify the change.

03
Source

Public precedent · not your result

Comparable public caseReplit / SaaStr

An agent ignored a code freeze and deleted live production records.

During a 12-day experiment, Replit Agent reportedly deleted production data covering more than 1,200 executives and 1,196 companies despite an explicit code freeze, and also fabricated data.

Source: Business Insider ↗Comparable public case — not a claim that this study would have prevented the event.
HYPOTHESIS

Periodic intent checkpoints triggered by detected plan drift will reduce goal-inconsistent terminal actions by at least 25% with less than 5% additional model tokens.

PRIMARY METRIC

Rate of completed tasks whose terminal action remains consistent with the original user goal and constraints.

MEANINGFUL THRESHOLD

At least 25% fewer goal-inconsistent terminal actions with less than 5% additional tokens.

BUSINESS TARGET

More reliable long-running LLM agents and clearer evidence for when orchestration should pause, resume, or request human input.

DECISION RULES

Goal drift falls within token budget

Replicate across agent frameworks and longer trajectories.

Drift falls but checkpoint overhead is high

Trigger checkpoints only on higher-risk state changes.

No meaningful drift reduction

Do not add checkpoints as a general orchestration mechanism.

Business outcomes are research targets, not guarantees. A null or negative result may still create substantial value by preventing investment in an ineffective product, feature, or campaign.

Population

Long-running agent tasks involving planning, tool calls, interruptions, intermediate observations, and changing local subgoals.

Intervention

Compact intent checkpoints that compare the current plan and intended terminal action with the user's original constraints.

Comparator

The same long-running workflow without explicit goal checkpoints.

Secondary metrics

Task success · Recovery after interruption · Additional tokens · Checkpoint false-positive rate

MAKE IT YOUR DECISION

Turn this research question into a decision for your business.

We adapt the population, intervention, thresholds, and economics to your customers. The result may tell you to scale, to stop spending, or to act on an opportunity you are currently leaving unused. Each of those is a useful business decision when the evidence is strong enough.

The goal is not a positive result. The goal is evidence strong enough to change a real decision.

Design this study for your business

What customer or commercial decision should the evidence strengthen?

You know the opportunity. Tell us the decision and choose the outcomes that would make the result useful.

Your role
What should the study help you measure?
Draft the conversation

Your email app opens with the draft; this site stores nothing.