Over a million developers have joined DZone.
{{announcement.body}}
{{announcement.title}}

Storing SysDig Streams with Apache NiFi

DZone's Guide to

Storing SysDig Streams with Apache NiFi

This is a tutorial on using Apache NiFi 1.0.0 to ingest, parse and store Linux system diagnostics and storing them in HDFS and file systems.

· Big Data Zone ·
Free Resource

The open source HPCC Systems platform is a proven, easy to use solution for managing data at scale. Visit our Easy Guide to learn more about this completely free platform, test drive some code in the online Playground, and get started today.

SysDig (GitHub) is an open source tool that allows for the exploration, analysis and troubleshooting of Linux systems and containers. It is well documented and very easy to install and use. It can be used for container and Linux system diagnostics, security analysis, monitoring and basic system information capture. Remember that SysDig can produce thousands of lines of messages and can continue doing so infinitely depending on the options.   

To install:To get started, follow the get started guide and try out the csysdig console.

 curl -s https://s3.amazonaws.com/download.draios.com/stable/install-sysdig | sudo bash 

There are a number of examples of using SysDig for getting specific system information, this could include processes, networking, files, errors, events and more.  

It is a very compact tool that can generate a ton of system and tracing information to use for current or predictive analysis of your system.   My use case was to capture a stream of this data every 15 minutes and store to Apache Phoenix on HBase for predictive analytics.

Image title

To use from the command line:

sysdig --help
sysdig version 0.11.0
Usage: sysdig [options] [-p <output_format>] [filter]

To start my flow in NiFi 1.00, I use the ExecuteProcess processor and wrap my SysDig command in a BASH script. This example returns ASCII text formatted as JSON and just gives me a 1-second snapshot.  This actually generates a ton of data.

 sysdig -A -j -M 1 --unbuffered 

This is an example of an Event returned by SysDig in JSON format:

"evt.cpu":6,
"evt.dir":">",
"evt.info":"fd=7(<f>/usr/lib64/python2.7/lib-dynload/_elementtree.so) ",
"evt.num":111138,
"evt.outputtime":1477313882635597873,
"evt.type":"fstat",
"proc.name":"python",
"thread.tid":14602}

Resources:

Managing data at scale doesn’t have to be hard. Find out how the completely free, open source HPCC Systems platform makes it easier to update, easier to program, easier to integrate data, and easier to manage clusters. Download and get started today.

Topics:
hadoop ,big data ,nifi ,logging ,sysdig ,devops

Opinions expressed by DZone contributors are their own.

{{ parent.title || parent.header.title}}

{{ parent.tldr }}

{{ parent.urlSource.name }}