The six practice areas below carry the escalated case from triage output to an analyst-ready intelligence brief. Each identifies what an agent can support, what evidence must remain visible, and where analyst judgment enters.
Map Agents to Intelligence Tasks
An agent earns its place in an intelligence workflow by handling repeatable, high-volume work while returning interpretive decisions to the analyst. The two ends of the workflow remain human-owned, where you define the intelligence requirement that scopes everything downstream, and an analyst approves the assessment before it is used.
In the first Refcard and its reference project, the triage workflow correlates a multi-step sequence that includes execution of a binary not present in the container image and escalates the case for human review. Here, the incident continues with an additional observable, the binary opening an outbound connection to cdn-sync.example.net.
Those signals create an incident-driven intelligence requirement: What is this binary, is the destination part of known malicious activity, and what other workloads may be exposed?
You own the requirement and its intended use because a poorly scoped question affects every stage that follows. An assessment prepared for internal incident response may require different sources, handling, and distribution than one intended for a shared advisory. An agent may draft or help refine the requirement, but it should not redefine it.
From there, the agent-supported workflow can collect matching indicators from approved sources and internal telemetry, normalize them into a consistent schema, enrich them with prior reporting and technique references, and propose possible candidate links. These outputs remain drafts, never findings.
The full path runs through eight stages, from the requirement that opens the case to the feedback that shapes the next cycle. Two review mechanisms run through it: light, stage-level supervision that scales with task risk and a single mandatory analyst gate at stage six before anything is used or shared.

Agent-Supported Intelligence Workflow
Stage six is the required approval point before dissemination. The analyst evaluates proposed correlations and any attribution against the evidence, source reliability, confidence, and original requirement. Unsupported or incomplete work returns to an earlier stage; approved work moves forward with a human-confirmed audience and handling designation.
Feedback then closes the loop by capturing what consumers used, what was noise, and what was missing, then feeding those findings into the next requirement.
Core Capabilities for Agent-Supported Intelligence
Five capability groups carry the workflow from raw feed to analyst-ready context. The table below groups them by function, mapping each stage in the figure to the task the agent supports, the output it produces, and the level of human review it requires.
The main distinction is between mechanical processing and interpretive work. Collection, deduplication, and normalization are high-volume tasks that may require only spot-checks on coverage and mapping accuracy. Enrichment and correlation carry greater risk. In the running example, the agent cross-references cdn-sync.example.net and the binary’s SHA-256 against external reporting, links related activity, and ties the result to internal asset and exposure context.
Any proposed campaign link remains a candidate until an analyst validates the linkage and its business-impact framing. Whether the feeds provide independent corroboration or echo one origin is addressed in the next section.
Confidence, uncertainty, and evidence-gap handling also belongs in the workflow as its own step. The agent can attach source-reliability and information-credibility ratings, carry the credibility assessment into STIX confidence, and flag what it could not establish. The analyst still owns the final confidence and interpretation.
Read the human review column as a gradient rather than a checkbox. Mechanical stages may need only a spot-check, while interpretive stages need substantive sign-off. This stage-level supervision is calibrated to risk and separate from the mandatory analyst gate at stage six in the workflow figure, which every item must clear regardless of how much review it received along the way.
Table: Core Capabilities for Agent-Supported Intelligence
| Capability Group |
Intelligence Task |
Output |
Human Review Needed |
| Collection and ingestion |
Pull from feeds, TAXII servers, OSINT, vendor reports, and internal telemetry; deduplicate; preserve source metadata and handling markings |
Consolidated, deduplicated items with provenance and original markings intact |
Light: spot-check coverage and source scope; confirm restricted sources are handled correctly |
| Normalization and structuring |
Parse heterogeneous inputs into a common schema (e.g., STIX objects); tag observed techniques against a shared framework (e.g., MITRE ATT&CK) |
Structured, machine-readable objects with consistent fields and technique references |
Light to moderate: validate ambiguous or novel mappings |
| Enrichment and correlation |
Cross-reference indicators, link related activity, and correlate external reporting with internal asset, exposure, and business context |
Enriched entities showing relationships, affected assets, and relevance to the environment |
Substantial: validate correlations and business-impact framing |
| Summarization and routing |
Draft analyst-ready summaries; route items by priority, topic, and sharing boundary (e.g., TLP label) |
Concise briefs with routing and assignment suggestions |
Moderate: confirm summary accuracy, routing, and markings before dissemination |
| Confidence, uncertainty, and evidence-gap handling |
Attach ratings for source reliability and information credibility (e.g., Admiralty A–F / 1–6; STIX confidence 0–100); flag single-source claims, stale data, and unresolved gaps |
Explicit confidence value, its basis, and a record of what could not be corroborated |
Substantial: determine final confidence and interpretation |
Evaluate Input Quality and Provenance
An agent’s brief is only as reliable as the inputs behind it. A summary can sound authoritative and still rest on a spoofed domain, a recycled block list, or three feeds echoing one origin. Weigh the inputs and their provenance before accepting the conclusion, assessing four properties separately: source reliability, information credibility, freshness, and independent corroboration.
Admiralty grading keeps reliability and credibility separate:
- Source reliability runs from
A, a source with a consistently reliable track record, to F, where reliability can’t be judged.
- Information credibility runs from
1, a claim confirmed by independent sources, to 6, when it can’t be evaluated at all.
The two combine into a rating like B2, “usually reliable, probably true.” The ratings stay separate because a trusted feed can still ship an unconfirmed indicator, while a new source can provide independently corroborated information. When intelligence is exchanged as STIX, the information-credibility judgment can be represented as a portable confidence value on a 0–100 scale.
Those judgments are difficult to verify unless the metadata travels with the data. For the dropped binary’s SHA-256 and the outbound host cdn-sync.example.net, each enrichment should identify who produced it, when it was created, how long it remains valid, its reliability rating, and its handling marking. Freshness carries real weight because attacker infrastructure rotates in hours; an IP flagged three weeks ago may now resolve to a clean CDN endpoint.
Three failure modes an agent may not reliably identify require additional review before trusting the enrichment:
- Feed manipulation or poisoning – Adversaries seed shared feeds with false indicators to force bad blocks or bury real activity. Although the risk has no single canonical definition, it maps conceptually to OWASP’s Data and Model Poisoning category.
- Stale indicators – An indicator past its validity window may continue to read as active when it isn’t.
- Circular reporting –
cdn-sync.example.net appearing in two feeds may look like corroboration, but if both trace back to one origin, it’s a single unverified claim counted twice. Check independence, not count.
A raw score doesn’t determine priority; internal context does. An A1 indicator tied to software you don’t run may rank below a medium-confidence indicator affecting an internet-facing critical asset. The agent attaches asset and exposure context, and the analyst decides what that context is worth.
Define Actionable, Reviewable Outputs
An agent-drafted brief tends to blur four things that must remain separate: the assessment, the observations behind it, unresolved hypotheses, and recommended actions. Label each one because a hypothesis presented as a finding may be acted on as fact. The finding carries the assessment, evidence pins each observation to a source, a hypothesis stays flagged as unconfirmed, and the recommendation only proposes. The final decision sits outside the brief entirely, belonging to the analyst at the mandatory review gate, not to anything the agent drafts.
When the brief lands in a shared queue, each audience looks for different information:
- An incident responder scans for affected assets and a block list.
- A detection engineer pulls the TTPs and IOCs worth turning into a rule.
- A team lead wants the one-line assessment and how far to trust it.
Analyst-ready means each audience lifts what they need without replaying the investigation.
A recommendation needs more than clear writing to be reviewable. Every claim should point to a specific observation and source, never a vague statement such as “analysis indicates.” It also has to be evidence-backed; the recommended action should name the supporting evidence so a reviewer can check the reasoning, not just the verdict. A good output also states its limits up front, including a single-sourced claim, a stale indicator, or a correlation inferred from timing rather than confirmed. That gaps field is often cut for space, but it’s where a confident summary can hide a thin case.
Below is a sample brief for the running example, the dropped binary calling out to an external host.
TLP:AMBER+STRICT |
|
Finding
Possible commodity malware staging on a production container: Image registry.example.com/app:1.4 ran a binary absent from its image layer, which then opened an outbound connection to cdn-sync.example.net. External reporting associates that host with malware-delivery infrastructure; no named-actor attribution. The outbound connection maps to MITRE ATT&CK T1071 (Application Layer Protocol). Evidence
- Runtime detector: A container from image
registry.example.com/app:1.4 in namespace prod executed a binary not present in the image, SHA-256 3f2b9c...e41 (placeholder), 2026-07-07 14:22Z [internal runtime telemetry, A1].
- Same process opened an outbound connection to
cdn-sync.example.net, resolving to 203.0.113.10 [internal runtime + DNS logs, A1].
- Destination host tagged as malware-staging infrastructure in two commercial feeds [external,
B3; single origin suspected, see Review Notes].
- Binary hash matches one public sandbox report [OSINT,
C3, single-sourced].
Source Confidence
Aggregate Admiralty B3 (usually reliable / possibly true); STIX confidence ~50, reflecting that corroboration is unconfirmed. Internal telemetry A1 (first-party). External host reputation B3. Hash-to-sandbox link C3, single-sourced.
Hypothesis
Unconfirmed: cdn-sync.example.net may be part of a broader malware-staging campaign rather than an isolated indicator. Flagged as unconfirmed because the two external feeds tagging this host may share one upstream origin, which would mean the reports aren’t independent corroboration. No named-actor or campaign attribution is asserted. The binary’s delivery mechanism is also unconfirmed: If it was staged over the network, that would align with MITRE ATT&CK T1105 (Ingress Tool Transfer), but no transfer event appears in current telemetry.
Affected Assets
The flagged container in namespace prod and its node; any workload running image registry.example.com/app:1.4. Potential blast radius: every replica pulled from that image tag across the cluster.
Recommended Actions
Analyst to confirm, then: isolate the workload with a NetworkPolicy denying egress and snapshot it for forensics before further interaction; block cdn-sync.example.net and 203.0.113.10 at egress; hunt the binary hash and destination host across 30-day egress logs; scan other workloads for image registry.example.com/app:1.4; add IOCs to the watchlist.
Suggested Priority
High
Review Notes
Hash-to-sandbox link is single-sourced (see Evidence, item 4). IP and domain use reserved documentation ranges; SHA-256 is a placeholder. Agent actions were read-only. |
These fields shorten the path from opening the brief to making the call. They separate the assessment from its evidence, flag what remains an open hypothesis, show how far to trust the reporting, identify the affected scope, keep recommended actions subject to human approval, and make unresolved gaps explicit.
Apply Human Review and Operating Guardrails
The control boundary comes first, where the agent drafts, and a human decides. Sort every output into two classes:
- Reversible internal work, like normalizing an indicator or drafting a summary, can move through the workflow automatically because a review still occurs before the output is used or shared.
- Anything operational or hard to undo, including publishing or blocking an indicator, asserting attribution, or sending a brief outside the team, requires explicit analyst sign-off.
Access should follow that boundary. Give the agent scoped, read-only access to threat data sources and no standing write or publish rights to detection, ticketing, or sharing platforms. Run it under a dedicated service identity so every action it takes is attributable and revocable. These restrictions help prevent excessive agency (OWASP LLM06) from turning a weak draft into a live action.
Before enrichment, minimize what leaves the organization. Strip or tokenize PII, victim identity, and internal-only context, and never ship restricted material to an external or hosted model that logs it. Handling labels should ride along here; the agent honors TLP markings and never downgrades one. It may propose an audience, but a person confirms before anything is shared. TLP ranges from TLP:RED for named recipients only to TLP:CLEAR for unrestricted sharing; AMBER+STRICT narrows AMBER to your organization only.
Treat agent output as untrusted until a reviewer validates it. A malicious report or sandbox artifact can contain instructions the model follows through prompt injection (OWASP LLM01), a generated indicator can be passed into a block or query unvalidated, and a fabricated or misattributed IOC can appear with confidence it hasn’t earned.
Reviewability also requires a paper trail. Each item should carry its provenance; the model version and prompting behind it; a review status of draft, reviewed, or approved; and the identity of the analyst who reviewed it, all in an append-only log. These are operating controls, not a governance program. The next section addresses how to measure whether they are effective over time.
Measure Actionability and Trustworthiness
Counting the indicators an agent ingests or the reports it drafts shows how busy the workflow is but not whether its outputs are useful. Volume can rise even when the agent is flooding the queue, so measurement should begin with the intelligence requirements the workflow is intended to support.
Ask whether an output advanced a standing requirement, reached the right consumer, and arrived while the decision it supported was still open. Consumer feedback can then show what share of the outputs were timely, relevant, and actionable, and how much of the work supported proactive intelligence instead of responding to an incident already underway. A report with richer context may take more effort to produce but still deliver greater value, so output count alone shouldn’t be treated as productivity.
Two additional signals evaluate the agent’s contribution. First, track how often a reviewer edits or discards a draft before it ships. Changes in the rate can expose over- or under-confidence as the workflow is tuned. Second, measure enrichment precision, the share of agent-proposed campaign links and context that survive analyst review.
Consumer judgements about accuracy, timeliness, and usefulness can reshape the next cycle’s intelligence requirements and inform adjustments to prompts, source weighting, and guardrails. Without that feedback, the workflow ossifies around the vanity metrics you intended to move beyond.