UM-Agent

The UM-Agent acts as a mediator. It extracts and prepares measurement data from various sources and sends it to monitoring or analytics systems. These sources may include log files, system parameters from Linux or Windows performance counters, the Windows EventLog, as well as outputs from any commands or programs that supply data from databases, APIs, etc.

Download Version 4.0.0
Source code on GitLab

Release Notes

Version 4.0.0 11.10.2026

Port of the Java version to the Go programming language. The configuration files remain largely compatible. The most important differences:

User Guide

Contents

  1. License
  2. Requirements
  3. Directory structure
  4. Installation
  5. Startup parameters
  6. Log files
  7. Configuration file
  8. Check group syswindows
  9. Check group eventlog
  10. Examples

License

The UM-Agent is released under the MIT License. It contains code from the following third parties. All trademarks and registered trademarks mentioned here are the property of their respective owners.

Go standard library
Copyright 2009 The Go Authors
License – BSD 3-Clause License https://go.dev/LICENSE

golang.org/x/sys
Copyright 2009 The Go Authors
License – BSD 3-Clause License https://cs.opensource.google/go/x/sys/+/master:LICENSE

Requirements

Directory structure

With an RPM installation, the program is located in /opt/umf/bin, the configuration in /etc/umf and the log files in /opt/umf/log.

Installation

On Linux with RPM (not finished)

The RPM package umagent-4.0.0-1.x86_64.rpm can be used on RHEL, Rocky Linux, AlmaLinux, CentOS Stream and other RPM-based distributions with systemd.

rpm -ivh umagent-4.0.0-1.x86_64.rpm
vi /etc/umf/umagent.xml
systemctl enable --now umagent

On Linux manually (not finished)

On any distribution it is enough to copy the program file.

mkdir -p /opt/umf/bin /opt/umf/log /etc/umf
cp linux/amd64/umagent /opt/umf/bin/umagent
chmod 755 /opt/umf/bin/umagent
cp umagent.xml /etc/umf/
cp umagent.service /usr/lib/systemd/system/
systemctl daemon-reload
systemctl enable --now umagent

On Windows

Extract UM-Agent_[version].zip

install.cmd 
sc.exe start umagent

The UM-Agent recognizes by itself that it was started as a service. The working directory is then the folder of the program file (bin), so the configuration is found at ..\cfg\umagent.xml and the log files end up in ..\log.

After the installation, the configuration files must be adjusted. You can use the supplied examples in the cfg folder as a basis. With umagent -list you can check the configuration in advance.

Startup parameters

Table 1 shows the startup parameters of the UM-Agent. The parameters may start with - or -- and may be abbreviated (-c, -cf, --cfg).

Table 1: Startup parameters for the UM-Agent

On Linux, the service is controlled with systemd (Table 2).

Table 2: Controlling the UM-Agent on Linux

The startup parameters of the service are in umagent.service (ExecStart). It is best to make changes with systemctl edit umagent.

Log files

In daemon mode, the agent records its work in two log files. If the program was started on the console, no log files are created and all messages appear on the console. The log files are rotated automatically when the date changes, and outdated log files are deleted automatically.

Log file umagent.log

The file umagent.log records the internal events in the following format:

Date Time Message-class Multi-check-ID Message

The message classes:

Checks that have the same source, e.g. the same log file or the same command, run as one unit called a multi-check. Right after the start, umagent.log records which checks a multi-check consists of, e.g.

2026.10.03 18:59:09 INFO MultiCheck-1 Start multi-check with members=1,2,3

Errors that occur before the log files are opened (e.g. configuration file not found or not valid XML) are written to the console. On Linux you can find them with journalctl -u umagent.

Log file mfdata.log

The file mfdata.log records the measurement results in the following format:

Date Time Parameter-number Parameter-name Value Unit

Example mfdata.log:

2026.10.03 15:28:20 3 Mem_Usage 99.07 %
2026.10.03 15:28:20 4 Swap_Usage 26.26 %
2026.10.03 15:28:20 5 Mem_Real_Free 6.98 %
2026.10.03 15:28:20 16 App_NET_IN 0.02 MB/s
2026.10.03 15:28:20 14 Alive_Status_OK 4.00 Msg/Min
2026.10.03 15:29:19 1 CPU_Total 10.30 %
2026.10.03 15:29:19 2 CPU_IOWait - %

A "-" as the value means: no measured value (NODATA).

Each measured value is sent to the Recipient depending on the sender (see global parameters). With UDP as a packet in the format object-code:parameter-number:value, e.g. SRV000000000001:1:10.30. With HTTP-POST and HTTP-GET as an HTTPS request, see Sender.

Configuration file

The configuration file defines which parameters the agent collects. The configuration file is a text file in XML format, the encoding is UTF-8 (ISO-8859-1 or Windows-1252 only with a matching XML declaration). Each parameter of a check can be given as an attribute <check parname="x"> or as an element <parname>x</parname>. Unknown attributes and elements are ignored and logged as a warning. In the following, mandatory parameters are marked [P] and optional parameters [O].

Global configuration parameters

The global configuration parameters control the overall work of the agent. Example:

<umagent>
    <common logdir="../log" keepdays="30" sender="UDP"/>
    <server host="monitor.example.com" port="9888" objcode="SRV000000000001"/>
    <include>
        <cfgfile path="../cfg/syst.xml"/>
        <cfgfile path="../cfg/appl.xml"/>
    </include>
    <check>
        ....
    </check>
</umagent>

Sender

The parameter sender in <common> defines how the measured values are transmitted to the Recipient. Only one sender is used at a time. All parameters of the configuration file (names and values such as sender, authtype) may be written in upper or lower case, e.g. HTTP-POST or http-post.

The checks pass their measured values to an internal queue, from which the sender reads and transmits them. A slow Recipient therefore does not slow down the checks. If the queue is full (4096 measured values), new measured values are lost; this is logged in the log file umagent.log.

Notes on HTTP-POST and HTTP-GET:

Example:

<common logdir="../log" keepdays="30" sender="HTTP-POST"/>
<server host="monitor.example.com" objcode="SRV000000000001"
        authtype="BASIC" authuser="agent" authpwd="secret"/>

Basic structure of a check

A check means the collection of the measured values of one parameter. There are five sources for the measurement data. Each source is called a group.

The groups syslinux and syswindows deliver measured values directly. The other three groups need an additional set of rules to extract the desired measured values from the strings. There are two types of rule sets:

All configuration parameters of a check can be divided into three areas:

Example of the general parameters of a check:

<check parname="." uom="." step="." parnum="." active="." objcode=".">
    <desc>...</desc>
    <group>...</group>
    <type>...</type>
    <calc operation=".."/>
    <default>N</default>
</check>

The conversion with <calc> is helpful in some cases, e.g. to convert the throughput of a network card from bytes/s to MB/s:

<calc operation="/ 1048576"/>

Or, if a check is executed in steps of 5 minutes, to convert the measurement result to a value per minute:

<calc operation="/ 5"/>

Check group logfile

The check group logfile processes log files in text format with ASCII or UTF-8 character encoding. Only lines added after the start of the agent are evaluated. A line is processed only when it is terminated by a line break. Rotating log files are recognized automatically if the step of the check is smaller than the step of the rotation. On Linux, if a log file was renamed or moved during the pause, the lines written to the old file during that pause are also taken into account. On Windows this is not the case: here the file is closed after each step so that the writing application can rotate it.

The type of the check specifies how measurement data is extracted. For the search, search rules are defined in <rule> elements. A search rule is either a regular expression in RE2 syntax (hereinafter Regexp) or an excerpt of a certain column (hereinafter Split). The prerequisite for the split is a log file in CSV format (Character Separated Values); the separator can be any single character. A check may contain several split rules and several Regexps. The search rules can be cascaded: each following rule searches the result of the previous rule. A Regexp returns the part in the first round bracket as its result. In 99.99% of cases, simple Regexps are sufficient.

Even in the last Regexp you must put the search result in round brackets. The string in the round brackets is interpreted as the final result (see Examples).

RE2 supports neither back-references (\1) nor lookahead/lookbehind expressions ((?=...), (?<!...)). With (?i) at the beginning, upper/lower case is ignored.

Type finde_digit

Using the search rules above, this check extracts the numbers from each new line of the log file or the output of a program and, depending on the setting, sends back the minimum, maximum, first value, last value, average or the sum of the numbers found. Integers and floating-point numbers are recognized as valid numbers:

E.g. "1.0020", "-0,23456", "+10000.12345686". Numbers with many digits are hard to read, so you can convert the result with the element <calc>. For example, converting bytes to MB:

<calc operation="/ 1048576"/>

If no new line appears in the log file or no match with a search rule was found, the default value is sent. If no default value is configured, "-" is sent; if the associated parameter on the Recipient has an evaluation rule, this is interpreted as NODATA. Example of a check of type finde_digit:

<check parname="." uom="." step="." parnum="." active=".">
    <group>logfile</group>
    <type>finde_digit</type>
    <path>/var/log/race.log</path>
    <default>0</default>
    <sendback>sum</sendback>
    <rule>
        <delimiter>,</delimiter>
        <column>5</column>
    </rule>
    <rule>
        <regexp>(^ \d{1,3}$)</regexp>
    </rule>
</check>
Type finde_string

The check searches for a certain text pattern in each new line of the log file and sends back the number of lines in which the search pattern occurs. A line is counted if the result of the last rule is not empty. Search rules work exactly as with the type finde_digit. If no new lines arrive in the log file or no match with the search pattern is found, 0 is reported. Example of a check of type finde_string:

<check parname="." uom="." step="." parnum="." active=".">
    <group>logfile</group>
    <type>finde_string</type>
    <path>../log/trace.log</path>
    <rule>
        <regexp>(\|ERROR\|)</regexp>
    </rule>
    <default>0</default>
</check>
Dynamic path

Some applications create log files whose path or file name changes over time, e.g. /var/log/app-20160929-10.log. The UM-Agent can handle such files if the path or file name contains the date, time or weekday. It is recognized automatically whether the path contains dynamic elements. The following dynamic elements are available:

Instead of the question mark, L or S must be used.

The example path on 29 September from 10:00 to 10:59:

/var/log/app-20160929-10.log

Corresponding path specification in the element <path>:

<path>/var/log/app-%LYYYY%%LMM1%%LDD%-%LHH24%.log</path>

The dynamic path contains the following elements:

Check group command

The check group command executes a command in every step and extracts the measured value from its output. By default the standard output (STDOUT) is evaluated, with <source>STDERR</source> the error output. The command is started directly without a shell; arguments containing spaces are put in quotation marks. For pipes, redirections or several commands in a row, call the shell yourself: on Linux sh -c "...", on Windows cmd /c "..." or powershell -NoProfile -Command "...".

If the command produces no output, the default value is sent. A return code other than 0 is logged as an error, but the output is still evaluated. If the command runs longer than <timeout>, it is aborted and the output received up to that point is evaluated. The two types finde_string and finde_digit are available for extracting the measurement data. Checks with the same command and the same step share one execution. Example of a check of group command with type finde_string:

<check parname="." uom="." step="." parnum="." active=".">
    <group>command</group>
    <type>finde_string</type>
    <execute>sh -c "ps -ef | grep apache"</execute>
    <timeout>2</timeout>
    <rule>
        <regexp>(Strang-1)</regexp>
    </rule>
</check>

Check group syslinux

The check group syslinux is available on Linux only and collects Linux-specific parameters from /proc, /sys and the command df. The element <type> specifies which statistic is needed. The following chapters list the available types. Checks of the same subgroup (cpu, mem, net, fs, bdev) with the same step are collected together. For CPU, network and block device, the values are calculated as the difference to the previous step; the first value therefore arrives after one step. An example can be found in /etc/umf/example-syslinux.xml.

CPU usage statistic

The CPU checks collect the statistics over all cores. Unit %. The following types are available:

RAM and swap statistic

The memory checks collect the statistics about RAM and swap. The following memory checks are available:

Network interface statistic

The network checks collect the statistics about any network interface. The checks need a logical network interface in the element <if>, e.g. <if>eth0</if>. If the interface does not exist at startup, the check is deactivated. The following network checks are available (values per second):

File system statistic

The file system checks determine the usage of the local file system under a mount point. Unit %. The check needs the path to the mount point in the element <path>. The check uses the Linux command "df", which is included in every Linux distribution. The following file system checks are available:

Block device statistic

The block device checks determine statistical data of a block device. A block device can be a hard disk, partition, RAID, LUN or multipath device, or a logical partition of one of these. This needs the path to the block device or mapper link in the element <path>. An important parameter of every block device is the sector size. It is used to calculate the data throughput and is read from the following file:

/sys/block/[Block-Device]/queue/hw_sector_size

The file system under "/sys" is not a real file system but a virtual file system for accessing kernel parameters at runtime. If the sector size cannot be read, the check returns no values and logs an error. The following block device checks are available:

A detailed description of the fields can be found in the kernel documentation (procfs-diskstats).

Check group syswindows

The check group syswindows is available on Windows only. It delivers system performance data from the Windows performance counters. The agent reads them directly through the Windows PDH interface. In the element <type> you specify a unique performance counter.

The names of the performance counters are always given in English, regardless of the language of the operating system. Names in the language of the operating system (e.g. "\Prozessor(_Total)\Prozessorzeit (%)") are not accepted. Examples:

How do I find the names of the performance counters?

The easiest way is the PowerShell script list-perfcounters.ps1, which is delivered together with the UM-Agent. It lists the performance counters with English names, as the agent expects them in the element <type>, on Windows in any language:

powershell -ExecutionPolicy Bypass -File list-perfcounters.ps1

Show the instances in an object, e.g. LogicalDisk

powershell -ExecutionPolicy Bypass -File list-perfcounters.ps1 -Object LogicalDisk -Instances
powershell -ExecutionPolicy Bypass -File list-perfcounters.ps1 -Object LogicalDisk -Instances | findstr /C:"Free Space"

Performance counters with * in their name must not be used!

Note on language: For some newer objects (e.g. "TCP/IP Performance Diagnostics", "SMB Server Shares", "QUIC Performance Diagnostics", "WMIPrvSE Health Status") there are no English names on a non-English Windows.

Alternatively, with typeperf.exe you can display the performance counters in the language of the operating system. On an English Windows the names can therefore be used directly:

Unknown performance counters are detected when the agent starts and deactivated, with an error message in the log file. With umagent -list you can check the names without starting the agent.

The values are collected at 5-second intervals. This produces several measurement results per step. Therefore the element <sendback> can be used. Default is "avg" (average). In some cases another function makes more sense, e.g. for the usage of a partition "last" is better suited. All syswindows checks with the same step are collected together. Example of a check of group syswindows with conversion from bytes to MBytes:

<check parname="DiscWriteC" uom="MB/s" step="60" parnum="7" active="1">
    <group>syswindows</group>
    <type>\LogicalDisk(C:)\Disk Write Bytes/sec</type>
    <calc operation="/ 1048576"/>
</check>

More examples in ..\cfg\example-syswindows.xml.

Check group eventlog

The check group eventlog is available on Windows only. The agent reads the events directly through the Windows event log interface. In every step, all new events that have occurred since the last step are read; in the first step all events since the start of the agent. Events from the time when the agent was not running are not read. Reading is limited to 10 seconds per step.

Each event is converted into one line; the fields are separated by "|":

Time|Level|EventID|Source|Text

The search rules are applied to this line. The two types finde_string and finde_digit are available for extracting the measurement data. The element <path> contains the name of the event log (channel). By default the following event logs exist on Windows: System, Application, Setup, Security – regardless of the language of the operating system. Channels such as "Microsoft-Windows-TaskScheduler/Operational" are also possible. If the channel does not exist at startup, the check is deactivated.

The example with type finde_string counts critical events in the System event log:

<check parname="EventSysCritical" uom="Error/Min" step="300" parnum="9" active="1">
    <desc>Check critical events in system eventlog.</desc>
    <group>eventlog</group>
    <type>finde_string</type>
    <path>System</path>
    <calc operation="/ 5"/>
    <rule>
        <regexp>^[^|]*\|(Critical)\|</regexp>
    </rule>
</check>

With split rules you can access a specific field. The example counts application crashes (EventID 1000 in the Application event log):

<rule>
    <delimiter>|</delimiter>
    <column>2</column>
</rule>
<rule>
    <regexp>^(1000)$</regexp>
</rule>

For testing purposes you can create an event yourself:

eventcreate /T ERROR /ID 158 /L APPLICATION /D "This is a test-event"

The test event is converted into the following line and then searched:

2026-10-03T21:45:24|Error|158|EventCreate|This is a test-event

Examples

The examples show the different ways of monitoring parameters in log files using the types finde_digit and finde_string.

finde_digit

For example, we search a log file for a certain pattern from which a number is to be extracted. Because the search pattern can occur in several lines, the average is to be sent to the Recipient. Here is an example line from the log file:

2012.05.20 20:23:23|INFO|Some-Modul...| ... Tried to connect XXX times|....||

This is achieved with the type finde_digit and a Regexp rule. The rule searches for the text "Tried to connect" followed by a number. The part in round brackets is interpreted as the number we are looking for. The check analyzes all new lines every 120 seconds. The average of all numbers found is sent to the Recipient with parameter number 1.

<check parname="ConnectTries" uom="-" step="120" parnum="1" active="1">
    <group>logfile</group>
    <type>finde_digit</type>
    <path>../log/random.log</path>
    <rule>
        <regexp>Tried to connect (\d+) times</regexp>
    </rule>
    <sendback>avg</sendback>
</check>

If the number you are looking for has decimal places (e.g. "123,2334"), you must extend the Regexp:

<regexp>Some result (\d+,\d+)</regexp>

finde_string

The check returns the number of lines from the log file in which the searched Regexp expression was found. The search rules work on the same principle, so here is just a list of various examples of Regexp expressions.

Table 3: Examples of regular expressions

Common errors


Copyright (C) 2012-2026 by Andrej Koslov. MIT License. · UMAgent Documentation · Software Version 4.0.0