DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Related

  • Offline Evaluation of RAG-Grounded Answers in LaunchDarkly AI Configs
  • A Practical Guide to Multi-Agent Swarms and Automated Evaluation for Content Analysis
  • ITBench, Part 2: ITBench User Experience – Democratizing AI Agent Evaluation
  • ITBench, Part 1: Next-Gen Benchmarking for IT Automation Evaluation

Trending

  • Gossips on Cryptography: Part 4
  • One Agent, Two Runtimes: Defining State Ownership Between Temporal and LangGraph
  • Kill the Worker, Keep the Research: Build a Recoverable LangGraph Agent on Temporal
  • The Silent Container Death: A TCP Dial That Never Times Out
  1. DZone
  2. Data Engineering
  3. AI/ML
  4. Resume the Evaluation, Not the Entire Batch: Build a Checkpoint-Aware AI Job Controller With Temporal

Resume the Evaluation, Not the Entire Batch: Build a Checkpoint-Aware AI Job Controller With Temporal

A checkpoint-aware Temporal controller resumes interrupted AI evaluations from saved progress instead of rerunning the entire batch.

By 
Akhil Madineni user avatar
Akhil Madineni
DZone Core CORE ·
Oct. 06, 26 · Tutorial
Likes (0)
Comment
Save
Tweet
Share
132 Views

Join the DZone community and get the full member experience.

Join For Free

Large AI evaluations rarely fail at convenient boundaries. A batch may contain tens of thousands of prompts, retrieval cases, tool-use scenarios, or judge-model comparisons, and each case can involve expensive network calls plus result persistence. Restarting the entire batch after a worker crash wastes inference spend and can change the meaning of the run when model outputs are nondeterministic. 

Temporal provides durable orchestration, but durability alone does not create application-level checkpoints. A reliable controller needs an explicit recovery boundary: completed evaluation cases stay completed, retries restart from a durable cursor, and the workflow remains small enough to replay efficiently.

Put the Retry Boundary Around Recoverable Work

The key design choice is the unit that gets retried. Temporal requires Workflow code to remain deterministic, while Activities are the place for external API calls and other nondeterministic work. That separation fits evaluation systems naturally: the Workflow owns lifecycle and policy, while an Activity calls the model, loads test cases, writes results, and advances progress. Temporal describes Activities as the failure-prone, side-effecting layer and supports retries and heartbeats for long-running work. 

A naive Workflow can schedule one Activity for the whole evaluation and rely only on an Activity retry. Without checkpoint logic, a retry re-enters the Activity from its method boundary, making case zero the recovery point. Scheduling one Activity per case avoids that problem but can create a large Workflow Event History when the dataset is large. A better middle ground is a checkpoint-aware Activity that processes a bounded slice of cases and records progress after each committed result.

The Workflow can keep retry behavior explicit rather than hiding it inside HTTP-client loops:

Java
 
ActivityOptions options = ActivityOptions.newBuilder()
    .setStartToCloseTimeout(Duration.ofHours(2))
    .setHeartbeatTimeout(Duration.ofSeconds(30))
    .setRetryOptions(RetryOptions.newBuilder()
        .setInitialInterval(Duration.ofSeconds(2))
        .setBackoffCoefficient(2.0)
        .setMaximumInterval(Duration.ofMinutes(1))
        .setMaximumAttempts(8)
        .build())
    .build();

EvaluationSummary summary =
    Workflow.newActivityStub(EvaluationActivities.class, options)
        .evaluate(spec);


StartToCloseTimeout bounds one Activity attempt, while the shorter heartbeat timeout lets Temporal notice a dead or disconnected worker much earlier. Temporal documents that a heartbeat timeout marks the Activity attempt as failed when heartbeats stop and allows another attempt to be scheduled according to the retry policy. 

Treat the Checkpoint as a Commit Record

A cursor such as nextIndex = 3840 is useful only when every result before that cursor is durably committed. The safe ordering is result write, checkpoint advance, then heartbeat. Reversing that order creates a lost-result window: a heartbeat could claim that case 3839 finished even though the corresponding result never reached durable storage. When result rows and checkpoints share a transactional database, both persistence operations can be committed atomically.

The evaluation identity also has to be immutable. A resumable run should pin the dataset snapshot, model identifier, inference parameters, prompt template, judge configuration, and scoring code version. Otherwise, a resumed process can silently mix results produced under different semantics. The checkpoint key should therefore be based on that immutable run identity rather than only on a human-readable batch name.

A compact Activity implementation can reconcile Temporal's last heartbeat with an external authoritative checkpoint:

Java
 
public EvaluationSummary evaluate(EvalSpec spec) {
    ActivityExecutionContext ctx = Activity.getExecutionContext();

    EvalCheckpoint heartbeat = ctx
        .getHeartbeatDetails(EvalCheckpoint.class)
        .orElse(EvalCheckpoint.start());

    EvalCheckpoint durable = checkpoints.load(spec.runId());
    int start = Math.max(heartbeat.nextIndex(), durable.nextIndex());

    for (int index = start; index < spec.caseCount(); index++) {
        EvalCase testCase = corpus.load(spec.datasetVersion(), index);
        EvalResult result = evaluator.run(spec.modelVersion(), testCase);

        results.upsert(spec.runId(), testCase.id(), result);
        EvalCheckpoint saved =
            checkpoints.advanceIfGreater(spec.runId(), index + 1, testCase.id());

        ctx.heartbeat(saved);
    }

    return results.summarize(spec.runId());
}


The external checkpoint is authoritative because Temporal heartbeats are a liveness and progress mechanism, not a transactional commit log for evaluation output. Temporal notes that heartbeats can be throttled by the Worker, and progress recorded immediately before a failure is available to the next attempt only if that heartbeat reached the Temporal Service before the Worker crashed.  The Math.max reconciliation tolerates a lagging heartbeat, while advanceIfGreater prevents an older attempt from moving the durable cursor backward.

Make Duplicate Execution Harmless

Checkpointing reduces repeated work but does not eliminate it. A worker can persist a model result and die before advancing the checkpoint, so the next attempt may execute the same case again. A timed-out attempt can also remain alive briefly from the perspective of an external system. The result store must therefore make duplicate case writes safe.

An idempotency key derived from runId and caseId provides a clean boundary. The persistence operation should insert once or perform a deterministic upsert instead of appending a second record. If the model provider supports request-level idempotency, the same stable key can be propagated there; otherwise, the local result store still prevents duplicate scoring records. The invariant is that repeating a case changes no already-committed state.

Checkpoint granularity then becomes an economic decision. Heartbeating after every case minimizes replay but increases heartbeat traffic. Heartbeating every small group reduces coordination overhead but increases the maximum repeated work after failure. Temporal Workers may throttle heartbeat delivery, so application correctness must never depend on every heartbeat being observed.  External progress commits can still happen at case granularity even when heartbeats are less frequent.

Heartbeats also provide a natural cancellation channel. Temporal delivers Activity cancellation through heartbeat calls, so a long evaluation loop that heartbeats regularly can stop promptly rather than continuing expensive inference after cancellation.  The checkpoint remains intact, making an operator-initiated retry or later continuation predictable.

Keep Long Evaluations Replayable

Checkpoint-aware Activities solve worker failure, but very large or continuously running evaluations can still grow Workflow history. Temporal's Continue-As-New mechanism starts a new Workflow Execution with fresh Event History while carrying forward the latest relevant state; the Workflow ID remains the same and the Run ID changes. Temporal recommends the mechanism when history becomes large or when long-lived executions need a fresh execution boundary. 

That mechanism works best with small Workflow state. A continuation should carry identifiers and cursors, not thousands of model responses. Large artifacts belong in a database or object store, with the Workflow retaining only stable references such as runId, dataset version, current shard, and checkpoint token. Temporal has separately warned against keeping excessive data in Workflow state and recommends external storage for large data. 

For a sharded evaluation, each Activity can own a deterministic range such as cases 20,000 through 24,999. After a shard completes, the Workflow records only the shard result reference and schedules the next range. A continuation boundary can be taken after a suitable number of shards:

Java
 
if (Workflow.getInfo().isContinueAsNewSuggested()) {
    ContinueState next = new ContinueState(
        spec.runId(), nextShard, aggregateRef);

    Workflow.continueAsNew(spec, next);
}


Temporal's Java documentation exposes isContinueAsNewSuggested() so application code can checkpoint state at a safe point before history limits become a concern.  This creates two complementary recovery layers: heartbeats and external checkpoints resume work inside an Activity, while Continue-As-New controls the lifetime and replay cost of the Workflow itself.

Conclusion

A reliable AI evaluation controller should treat recovery as a data-consistency problem rather than a generic retry problem. Temporal supplies durable execution, retry policies, heartbeat-based failure detection, cancellation delivery, and fresh-history continuation, but the application still defines what “completed” means. 

Persisting each case idempotently, advancing a monotonic external checkpoint only after that persistence succeeds, and heartbeating the resulting cursor turns a failed worker into a small replay window instead of a full-batch restart. Keeping immutable evaluation semantics attached to the run and using Continue-As-New for long histories preserves both correctness and operational efficiency. The result is a controller that resumes the evaluation at the last trustworthy commit point rather than paying again for work that has already finished.

AI Evaluation Checkpoint (pinball)

Opinions expressed by DZone contributors are their own.

Related

  • Offline Evaluation of RAG-Grounded Answers in LaunchDarkly AI Configs
  • A Practical Guide to Multi-Agent Swarms and Automated Evaluation for Content Analysis
  • ITBench, Part 2: ITBench User Experience – Democratizing AI Agent Evaluation
  • ITBench, Part 1: Next-Gen Benchmarking for IT Automation Evaluation

Partner Resources

×

Comments

The likes didn't load as expected. Please refresh the page and try again.

  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook