DZone
Thanks for visiting DZone today,
Edit Profile
  • Manage Email Subscriptions
  • How to Post to DZone
  • Article Submission Guidelines
Sign Out View Profile
  • Post an Article
  • Manage My Drafts
Newsletter
Log In / Join
Refcards Trend Reports
Events Video Library
Refcards
Trend Reports

Events

View Events Video Library

Related

  • Building a Secure MCP Server for File Processing: Auth, Rate Limiting, and Idempotency
  • How to Build a Production-Ready iOS App With AI-Generated Code
  • Your Application Has an Unindexed Attack Surface. Do You Know What’s in It?
  • Why Continuous Application Security Testing Is No Longer Optional

Trending

  • A Senior Engineer’s Guide to Foundry IQ, MCP, and the OpenAI Agents SDK
  • Building Agentic RAG, Step by Step: From Static Retrieval to Reasoning Pipelines
  • How to Perform Response Verification in REST-Assured Java for API Testing: Part 2
  • How to Correctly Implement ‘Sneaky Throws’ in Java
  1. DZone
  2. Software Design and Architecture
  3. Security
  4. Detection and Response Did Its Job. Now Someone Has to Actually Fix It.

Detection and Response Did Its Job. Now Someone Has to Actually Fix It.

Here’s how attack evidence, ownership mapping, root cause tracing, developer routing, and retesting turn a contained incident into an actual fix.

By 
Philip Piletic user avatar
Philip Piletic
DZone Core CORE ·
Sep. 29, 26 · Analysis
Likes (0)
Comment
Save
Tweet
Share
127 Views

Join the DZone community and get the full member experience.

Join For Free

A host is isolated. A malicious process is killed. A compromised credential is revoked before an attacker can use it again. Detection and response worked exactly as intended, automatically, correctly, and fast.

What follows is usually less tidy. The attacker may be gone, but the opening they used can still be sitting there waiting for the next attempt. Veracode’s State of Software Security Report gives some sense of the problem’s scale. Critical security debt was up 20% year over year, and high-risk vulnerabilities, those considered both severe and highly exploitable, climbed 36%. Detection has made progress, but finding a problem quickly and implementing fixes are clearly not the same thing.

A cybersecurity platform that’s actually good at threat detection and incident response has to do more than contain the immediate event. The incident still has to be reconstructed, connected to the team responsible for the affected service, and followed back to whatever condition made the attack possible. Only then can developers or IT managers make the change and return to the environment to see whether it worked.

Reconstructing What Actually Happened

An isolated host does not come with a narrative attached. Neither does a killed process. What the security team initially has is an event, and before someone can fix the underlying problem, they need to reconstruct the sequence that produced it.

Sysdig offers a useful example of what that evidence can look like. Its Falco-based real-time detection can sit alongside response actions such as killing a process, pausing a container, or quarantining a file. But stopping the activity is only part of the process. The surrounding runtime evidence can show an investigator what was happening on the system when the alert fired, rather than leaving the team to piece together the incident after the fact.

That evidence can include processes being executed, network connections, changes to files, and the lineage between processes. On Linux hosts, Kubernetes nodes, and VMs, Sysdig can collect system-call data using eBPF-based drivers, with kernel modules also supported. Serverless environments require instrumentation appropriate to that execution model, while Windows uses Event Tracing for Windows rather than Linux eBPF for kernel-level workload visibility.

What all of this data has in common is that it captures actual behavior, not just another alert. It provides evidence of what software did in production, evidence that becomes the raw material for figuring out what really happened during an incident.

That evidence is what runtime security is actually for. Instead of merely producing another alert, it’s most useful when scans reveal what software did in production once something has already gone wrong. Fixing what made that possible is a separate job, which is exactly why runtime and build-time security have to work together rather than standing in for each other.

Finding Whose Service This Actually Is

Reconstructing the attack is one problem. Figuring out whose problem it is can be another. A workload running in production carries plenty of technical information, but it does not necessarily tell the responder which repository produced it, which engineering team is responsible for it, or who is in a position to change it without breaking something else.

That ownership gap becomes particularly difficult in cloud-native environments. A compromised container may belong to a service assembled from several repositories, deployed through shared infrastructure code, and operated by a platform team that did not write the vulnerable application. A credential may technically belong to one cloud account while being consumed by workloads owned elsewhere.

Most organizations already have clues scattered across their environment. A service catalog may name an owner, CODEOWNERS may point to a team, repository metadata may identify maintainers, and infrastructure tags or deployment history can help connect what is running back to where it came from. None of those records is particularly useful, though, if it describes an organization that existed six months ago rather than the one responding to the incident today.

Ultimately, the security team needs to get from “this workload was compromised, and our automation has contained it” to the much more useful, “we know which team can remediate the issue that made this situation possible.”

Tracing Why It Was Actually Exploitable

At this stage, the responder knows the shape of the incident and has a reasonable idea of who should own the fix. What is still missing is the cause. The question shifts from what the attacker did to what was present in the environment that let those actions succeed.

The process that was killed may only be the last link in a much longer chain. Perhaps an old dependency made it into the container image. Maybe a role had permissions it never needed, a service was reachable from somewhere it should not have been, or an infrastructure configuration quietly exposed a path into the workload. Unless that earlier condition changes, stopping the process deals with the incident without really dealing with its cause.

This is the part of the workflow Wiz’s Green Agent is designed to address, the stage most detection and response tooling stops short of. Once a runtime finding has been detected and contained, Green Agent picks up from there, analyzing the confirmed finding in the context of the environment, tracing the issue back to its root cause, identifying ownership, and generating environment-specific remediation guidance for the developer or owner positioned to make the change.

The Security Graph provides supporting context across areas such as identities, network exposure, cloud resources, code, and data sensitivity, helping distinguish the immediate runtime event from the condition that made it exploitable, so detection and response extend past containment into an actual fix.

Tracing a runtime event back to the code that caused it isn’t a new problem. Application security teams have been working on closing the gap between what SAST catches in code and what DAST catches at runtime for years. A finding is only actionable once it’s connected back to its source. Incident remediation is running into the same wall now, just with an attacker involved instead of a scanner.

Getting the Fix Into the Developer’s Workflow

Root-cause analysis can still fail operationally if its result lives inside a security console that the responsible developer rarely opens. The handoff therefore matters almost as much as the diagnosis. Security teams need to decide which findings require human intervention and then put those findings into systems where engineering work is already managed.

Palo Alto Cortex XSIAM handles the triage side of that decision. Embedded automation enriches alerts and closes low-risk cases before they ever reach an analyst’s queue, leaving higher-value cases for actual investigation and response.

The developer who has to make the change probably isn’t working out of a security console at all. Their day is more likely to revolve around a repository, an issue tracker, a pull request, an IDE, and team messages. That reality colors what a useful handoff looks like. Sending another alert is not enough. The developer needs enough of the incident’s story to understand why the change is being requested and enough technical context to know where to start.

The lesson is familiar from application testing. Making DAST findings actionable requires narrowing the distance between discovering a vulnerability and getting it into a form developers can realistically act on. Runtime incidents create much the same handoff problem, only with an attacker potentially having demonstrated the consequence already.

Confirming the Fix Actually Held

The last stage is easy to treat as administrative cleanup, but it is part of remediation itself. A developer changes a dependency, tightens a role, modifies a configuration, or removes an exposure. The original condition then needs to be tested again.

A clean retest is useful evidence that the remediation changed the behavior security observed. It is not absolute proof. An endpoint may have moved, an access path may have changed, or another control may now be masking the original condition without eliminating its cause.

The same caution applies to incident response. Reconnecting a host or restoring a quarantined workload does not demonstrate that the weakness behind the incident is gone. Without retesting, the organization has confirmed that containment can be reversed, not that remediation succeeded.

Conclusion

Automated detection and response is extremely effective at handling emergencies. It can interrupt malicious activity faster than a human analyst could reasonably investigate and act.

What Veracode’s numbers suggest, however, is that stopping the immediate event has not solved the industry’s bigger challenges. The work that follows is slower and crosses more boundaries, from forensic evidence to service ownership, engineering changes, and finally another look at the environment to make sure the original weakness is no longer there.

That full path is what separates a genuinely complete approach to detection and response from an approach that only detects and contains. A response system can stop the attacker and give the organization breathing room, but somebody still has to use what the incident revealed to remove the condition behind it. Then the organization has to check its work. Isolation buys that opportunity, and remediation is what makes use of it.

security

Opinions expressed by DZone contributors are their own.

Related

  • Building a Secure MCP Server for File Processing: Auth, Rate Limiting, and Idempotency
  • How to Build a Production-Ready iOS App With AI-Generated Code
  • Your Application Has an Unindexed Attack Surface. Do You Know What’s in It?
  • Why Continuous Application Security Testing Is No Longer Optional

Partner Resources

×

Comments

The likes didn't load as expected. Please refresh the page and try again.

  • RSS
  • X
  • Facebook

ABOUT US

  • About DZone
  • Support and feedback
  • Community research

ADVERTISE

  • Advertise with DZone

CONTRIBUTE ON DZONE

  • Article Submission Guidelines
  • Become a Contributor
  • Core Program
  • Visit the Writers' Zone

LEGAL

  • Terms of Service
  • Privacy Policy

CONTACT US

  • 3343 Perimeter Hill Drive
  • Suite 215
  • Nashville, TN 37211
  • [email protected]

Let's be friends:

  • RSS
  • X
  • Facebook