Linux Audit and File Integrity Monitoring (auditd)

This guide covers file integrity monitoring on a Linux host that is already reporting to Secure60: the watches to add, the detection rules to deploy, the event volume to manage, and the handling of patch windows. The audit baseline itself is installed by the Linux Server - Integration Guide.

Prerequisite

auditd installation, the syslog plugin, the Secure60 audit rule baseline, forwarding and verification are covered by the Linux Server - Integration Guide. Nothing on this page takes effect until that setup is complete.

Overview

Secure60 file integrity monitoring on Linux uses the audit subsystem already present in the operating system. auditd reports every write and attribute change to a watched path as it happens, with the process and the login uid that made it. The record reaches Secure60 over the syslog path opened for the host’s other logs, and the managed File Integrity rule group evaluates it. No agent is installed and no baseline snapshot is maintained on the host.

FILE INTEGRITY MONITORING LIFECYCLE Audit baseline 57-rule ruleset syslog forwarding File watches own paths, own keys 60-fim-local.rules Detection rules managed rule pack deployed per project Maintenance window MAINTENANCE group in before, out after Review residue per window threats stay armed tune watches, execve rules and rate limits after each review

The lifecycle has five parts, each covered below:

  1. Audit baseline — the 57-rule Secure60 ruleset and syslog forwarding, from the Linux Server - Integration Guide.
  2. File watches — the paths specific to this estate, added under s60_fim_* keys (Custom file integrity watches).
  3. Detection rules — the managed File Integrity rule group, deployed to the project (Deploying the File Integrity rules).
  4. Maintenance windows — hosts placed in the MAINTENANCE entity group for the duration of a patch (Patching and maintenance windows).
  5. Review — the events that remain after a window, and the tuning of watches, execve rules and rate limits (Managing the volume, Rate limits).

Baseline coverage

The baseline covers process execution and privilege escalation, identity and group management, sudo, SSH and PAM configuration, cron, systemd and profile persistence, kernel module loading, dynamic loader configuration, system time, network and name resolution, mandatory access control, audit configuration tamper, login and session records, process inspection, and mounts. Each group carries its own s60_* rule key, and the managed rules are written against those keys. This page adds to that baseline; it does not replace any of it.

Custom file integrity watches

A watch is a path, a permission mask, and a key:

-w /etc/nginx/          -p wa -k s60_fim_web
-w /opt/payments/config -p wa -k s60_fim_app
-w /usr/local/bin/      -p wa -k s60_fim_bin

w records writes and a records attribute changes. r (read) on a busy directory produces an event for every file access and is rarely warranted.

Configuration and binary directories are the paths worth watching. A change under /var/lib/mysql is routine; a change under /usr/local/bin is not.

The s60_fim_ key prefix matters: the managed rule File Integrity - Monitored Path Changed fires on any key with that prefix, so every watch written this way is detected without Secure60 holding the path list.

Watches belong in their own file, for example /etc/audit/rules.d/60-fim-local.rules, so that a later update to the Secure60 baseline does not overwrite them. Then:

sudo augenrules --load
sudo auditctl -l
One bad line stops every rule below it

augenrules --load halts at the first line it cannot load. Rules above that point are live; everything below is silently absent, including rules from other files that merge after it.

A watch resolves its path when the rule loads, and fails if the parent directory is missing. The file itself need not exist: /etc/ld.so.preload loads on a host that has never had one, which is what makes its creation detectable.

Because the watches file sorts after the Secure60 baseline, a bad path in it cannot break the baseline, but it drops every watch below it. After any change:

sudo augenrules --load          # any "There was an error in line N" is a hard failure
sudo auditctl -l                # confirm the LAST rule you wrote is present

Which record carries the filename

For a file integrity event, the SYSCALL record identifies who did it and with what, but the path is not on it. The filename lives on the accompanying PATH records, alongside file_inode, file_mode and file_owner_uid, and the Linux Platform parser combines every record that shares an event_audit_id into one event. One syscall produces several PATH records: the parent directory always comes first, and a rename carries both the old name and the new one. The parser reads nametype on each record to work out which name is which:

Field What it holds
event_target_filename The file the operation was about: the file created, written, deleted or executed, or the new name after a rename or link
event_source_filename On a rename, the name the file had before; on a hard link, the file the new name points at. Absent otherwise
file_directory The directory event_target_filename sits in
event_directory The working directory the process ran from, when the CWD record is collected. Relative names are resolved against it
event_nametype What the target was to the syscall: CREATE, DELETE, NORMAL (an existing file) or UNKNOWN (the name could not be resolved, typically a failed call)

mv /etc/cron.d/job.tmp /etc/cron.d/job is therefore one event with operation = 'file-rename', event_target_filename = '/etc/cron.d/job' and event_source_filename = '/etc/cron.d/job.tmp'. This is the shape every package manager and editor produces, since they write a temporary file and rename it into place.

Deploying the File Integrity rules

Audit events are evaluated by the managed rule group Secure60 - Managed Rules - File Integrity. It is published to every organisation and runs nothing until it is deployed to a project. Deployment, the update policy and the per-project controls are described in Rules and Rule Groups; the AUTO update policy is recommended for this pack, because the process attribution lists in the rules (package managers, relabel tools, network and time daemons) are maintained by Secure60.

The pack contains eleven rules. Three raise a threat; the rest score the host and the user entity so that a host accumulating unexplained change surfaces through entity analytics.

Rule Fires on Outcome
Audit Configuration Changed s60_auditcfg — /etc/audit/, the rules themselves Threat, once per host
Dynamic Loader Configuration Changed s60_loader — ld.so.conf, ld.so.conf.d/, ld.so.preload Threat, once per host
Kernel Module Loaded or Removed s60_modules — module syscalls and modprobe.d/ Threat, once per host
Audit Reporting Stopped No audit record from a host for six hours Threat, once per host
Track Audit Reporting Any audit record Entity record, no score
Monitored Path Changed Any s60_fim_* key Scored signal
Identity or Privilege File Changed s60_identity, s60_privcfg Scored signal
Authentication Configuration Changed s60_authcfg — sshd and PAM configuration Scored signal
Persistence Location Changed s60_persist — cron, systemd, shell profiles Scored signal
Login or Session Record Changed s60_logins, s60_session Scored signal
System Configuration Changed s60_netcfg, s60_mac, s60_time, s60_mount Scored signal

The rules also recognise the CIS-style key names that estates commonly already run (identity, logins, time-change, MAC-policy and similar), so a host carrying an existing third-party ruleset is covered without renaming its keys.

Custom rules and the MAINTENANCE group

Every scored rule in the pack carries the condition !isEntityGroup('MAINTENANCE'), which is what makes the maintenance window described below work. The three threat rules do not carry it, by design.

A rule written by the customer against its own watch keys does not inherit that condition. It must be included in the rule’s query, or the rule fires throughout every patch window:

event_audit_key='s60_fim_app' AND operation!='file-access' AND user_name!='' AND !isEntityGroup('MAINTENANCE')

user_name!='' restricts the rule to changes the audit subsystem attributes to a login uid, which excludes the configuration-management and monitoring agents that rewrite files on a schedule. See Query Syntax for isEntityGroup.

Managing the volume

One rule group in the baseline generates meaningful volume: process execution. Measured on a steady-state Rocky 9 host, every other group produced zero records in sixty seconds. The recipes in this section therefore narrow execve.

Each process launch produces a SYSCALL record and an EXECVE record as a pair, plus PATH and CWD records for the paths involved.

PROCTITLE exclusion

The baseline excludes PROCTITLE, which repeats, hex-encoded, the command line the EXECVE record already carries. Across the Secure60 fleet its record count tracks SYSCALL almost exactly, so it is one record in seven with no detection value.

-a always,exclude -F msgtype=PROCTITLE
Excluding PATH removes the filename from every file event

-F msgtype=PATH is widely recommended alongside PROCTITLE, and on a host doing execution auditing only it costs nothing.

On a host running file watches it removes the filename from every event they produce. A write to /etc/shadow still produces a SYSCALL record with user, process, syscall and outcome, and no record names the file. CWD carries the working directory needed to resolve relative paths, and excluding it has the same effect on those.

Where PATH volume is the concern, the execve rules are the source of it; file watches generate few PATH records. Narrow the execve rules as below.

Narrowing the execve rules

In order of benefit per unit of effort.

Drop the 32-bit rule on 64-bit-only hosts. If nothing on the host runs 32-bit binaries, arch=b32 produces nothing:

-a always,exit -F arch=b64 -S execve -k s60_exec

Exclude execs with no logged-in user behind them. Daemons and system services run with an unset audit UID. Excluding them keeps every human-initiated command and removes most of the background churn:

-a never,exit -F arch=b64 -S execve -F auid=unset

On older audit versions, write -F auid=4294967295 instead of auid=unset.

This exclusion also hides the activity a webshell produces: a compromised web server executing commands as www-data has no audit UID either. The s60_privesc and s60_modules rules still fire. On internet-facing hosts, excluding named binaries is the safer option.

Exclude specific noisy binaries. Monitoring agents, backup jobs, and health checks are usually the top few entries:

-a never,exit -F arch=b64 -S execve -F exe=/usr/bin/node_exporter
-a never,exit -F arch=b64 -S execve -F exe=/usr/lib/nagios/plugins/check_disk

Audit privileged execution only. The narrowest useful option records everything running as root and nothing else. It gives up all visibility of unprivileged attacker activity, so the s60_privesc rules stay alongside it:

-a always,exit -F arch=b64 -S execve -F euid=0 -k s60_exec_root
Order decides which rules can fire at all

For a given syscall the kernel walks the rule list and stops at the first match. A never rule placed below the always rule it is meant to suppress never runs, and neither does a narrow always rule below a broad one covering the same syscall. There is no error, and auditctl -l shows nothing unusual.

never rules go above the always rule they modify. This is also why s60_privesc sits above s60_exec in the baseline.

Additional syscall rules

These carry detection value but are high-volume on a busy host, so they are added deliberately rather than as part of the baseline:

Syscalls to audit Catches Volume
chmod, fchmod, fchmodat, chown, fchown, fchownat, setxattr, lsetxattr, fsetxattr Permission and ownership tampering High
unlink, unlinkat, rename, renameat Bulk file deletion and log destruction High
open, openat, openat2, truncate, ftruncate (with -F exit=-EACCES, then again with -F exit=-EPERM) Failed access attempts — reconnaissance Very high

Each needs the usual -a always,exit -F arch=b64 prefix, the syscalls comma-separated after -S, and a -k key.

Measuring the effect

In Search, query for the host and use Analytics to group by host_name over 24 hours, then compare per-host event counts before and after a rule change. A step change appears within minutes of augenrules --load. See Search.

Sizing an estate

Process execution auditing on a busy server is commonly the dominant share of that host’s ingest. When estimating volume for a new deployment, enable the baseline on one representative host, measure a full day, and multiply that figure across the fleet rather than using a general estimate.

Patching and maintenance windows

An OS patch cycle produces more file integrity events than any other routine operation on a watched host. A single dnf upgrade writes to /etc/pam.d/, /etc/ld.so.conf.d/, /usr/lib/systemd/system/, /etc/profile.d/ and, through the SELinux policy package, hundreds of files under /etc/selinux/; the reboot that follows loads every kernel module the hardware needs. Each of those is a legitimate file integrity event, and without context each would score the host and the administrator who ran the update.

The managed File Integrity rules handle this in two layers: attribution by the writing process, applied by the rules themselves, and a maintenance-window group, set by the customer.

Attribution by writing process

The audit record for a change carries the process that made it and the login uid it was made under, and the rules use both:

The MAINTENANCE group

For a planned change window, the hosts are placed in an entity group named MAINTENANCE before the work starts and removed when it finishes. Every scored File Integrity rule excludes members of that group for as long as they are in it. The three threat rules — audit configuration, dynamic loader, kernel module — stay armed inside the window by design; those changes warrant review during a patch.

The group is not created by Secure60; each organisation creates its own, once. The group name is matched exactly. If the organisation has no group of that name, the rules behave as though the exclusion is absent.

From the portal: open Entities → Groups, create a group named MAINTENANCE once, then open it and assign the hosts. Unassign them when the window closes.

From a script, for example a pre-task and post-task around a patch run. Group membership is a link on the host entity, and a write replaces all links of that type, so the host’s current groups are read first and added to:

API=https://api.secure60.io/admin/1.0
ORG=<organisation_id>
HOST=<host_name>          # exactly as the host reports it in host_name
AUTH="Authorization: Bearer $SECURE60_TOKEN"

GROUP=$(curl -s -H "$AUTH" "$API/entity?organisation_id=$ORG&type=group&name=MAINTENANCE" | jq -r '.result[0].id')
HOSTID=$(curl -s -H "$AUTH" "$API/entity?organisation_id=$ORG&type=host&value=$HOST" | jq -r '.result[0].id')
CURRENT=$(curl -s -H "$AUTH" "$API/entity/$HOSTID?organisation_id=$ORG&include_links=Y&links_type=host_in_group" \
          | jq -c '[.result.links[]?.targets[]?.id]')

# Enter the window: add MAINTENANCE to the host's groups
TARGETS=$(jq -cn --argjson c "$CURRENT" --argjson g "$GROUP" '$c + [$g] | unique')
curl -s -X PUT -H "$AUTH" -H 'Content-Type: application/json' "$API/entity/$HOSTID" \
     -d "{\"organisation_id\":$ORG,\"links\":[{\"link_type\":\"host_in_group\",\"targets\":$TARGETS}]}"

# Leave the window: remove it again
TARGETS=$(jq -cn --argjson c "$CURRENT" --argjson g "$GROUP" '$c - [$g]')
curl -s -X PUT -H "$AUTH" -H 'Content-Type: application/json' "$API/entity/$HOSTID" \
     -d "{\"organisation_id\":$ORG,\"links\":[{\"link_type\":\"host_in_group\",\"targets\":$TARGETS}]}"

A host that has never been registered as an entity has no HOSTID; it is created first with POST $API/entity and {"type":"host","value":"<host_name>","name":"<host_name>","definition":{},"organisation_id":<organisation_id>}.

From Ansible, the same four calls become pre_tasks and post_tasks around the patch play, so the window opens and closes with the run and a failed play leaves the host in the group for review. S60_TOKEN and S60_ORG come from the environment; the pause is the membership cache described below.

- hosts: linux_servers
  gather_facts: false
  vars:
    s60_api: https://api.secure60.io/admin/1.0
    s60_org: "{{ lookup('env', 'S60_ORG') }}"
    s60_auth: { Authorization: "Bearer {{ lookup('env', 'S60_TOKEN') }}" }

  pre_tasks:
    - name: Resolve the MAINTENANCE group and this host's entity
      delegate_to: localhost
      uri:
        url: "{{ s60_api }}/entity?organisation_id={{ s60_org }}&type={{ item.type }}&{{ item.by }}={{ item.value | urlencode }}"
        headers: "{{ s60_auth }}"
      loop:
        - { type: group, by: name,  value: MAINTENANCE }
        - { type: host,  by: value, value: "{{ inventory_hostname }}" }
      register: s60_lookup

    - name: Keep the ids
      set_fact:
        s60_group_id: "{{ s60_lookup.results[0].json.result[0].id }}"
        s60_entity_id: "{{ s60_lookup.results[1].json.result[0].id }}"

    - name: Read the host's current groups (a write replaces all host_in_group links)
      delegate_to: localhost
      uri:
        url: "{{ s60_api }}/entity/{{ s60_entity_id }}?organisation_id={{ s60_org }}&include_links=Y&links_type=host_in_group"
        headers: "{{ s60_auth }}"
      register: s60_entity

    - name: Keep the current group ids
      set_fact:
        s60_groups: "{{ (s60_entity.json.result.links | default([], true)) | map(attribute='targets') | flatten | map(attribute='id') | map('string') | list }}"

    - name: Enter the maintenance window
      delegate_to: localhost
      uri:
        url: "{{ s60_api }}/entity/{{ s60_entity_id }}"
        method: PUT
        headers: "{{ s60_auth }}"
        body_format: json
        body:
          organisation_id: "{{ s60_org | int }}"
          links:
            - link_type: host_in_group
              targets: "{{ (s60_groups + [s60_group_id]) | unique }}"

    - name: Wait for the membership cache
      pause:
        seconds: 180
      run_once: true

  tasks:
    # The patch work goes here — dnf upgrade, reboot, health check.
    - name: Patch
      debug:
        msg: "patching {{ inventory_hostname }}"

  post_tasks:
    - name: Leave the maintenance window
      delegate_to: localhost
      uri:
        url: "{{ s60_api }}/entity/{{ s60_entity_id }}"
        method: PUT
        headers: "{{ s60_auth }}"
        body_format: json
        body:
          organisation_id: "{{ s60_org | int }}"
          links:
            - link_type: host_in_group
              targets: "{{ s60_groups | difference([s60_group_id]) }}"

inventory_hostname must be the name the host reports in host_name; an inventory that uses short names or IPs needs a host variable in its place.

Membership cache: three minutes

Group membership reaches the query layer through a cache that refreshes every three minutes. A host added at the same moment the first package is written scores for the first minutes of the window. Assigning the hosts three minutes before the change starts avoids it.

Review after the window

What fires during a patch is, by construction, what was not written by the package manager: a shell writing to a watched path, a change under a login uid to a persistence location, a new file in /etc/ld.so.conf.d/ that no package put there. The search product='auditd' AND isField('event_audit_key') AND user_name!='' over the window returns that list, with the process and user on each line. It is the review record an assessor expects after a change.

auditd disk limits

Execution auditing writes a large local audit.log as well as forwarding. The defaults in /etc/audit/auditd.conf decide what happens when the disk runs out, and some of them take the host down:

Setting Recommended Why
max_log_file_action ROTATE KEEP_LOGS fills the filesystem
space_left_action SYSLOG Warns while there is still time to act
admin_space_left_action SYSLOG or SUSPEND HALT powers the host off
disk_full_action SUSPEND HALT powers the host off

max_log_file × num_logs is the ceiling on local audit retention. The difference between 8 × 5 and 50 × 30 is 40 MB against 1.5 GB per host.

The kernel-side failure mode is set in the rules file, not auditd.conf. -f 1 prints a warning when the audit backlog overflows; -f 2 panics the kernel. Most distributions ship -f 1 in 10-base-config.rules; confirm before enabling high-rate auditing.

The backlog and loss counters are the ongoing health check:

sudo auditctl -s

A rising lost counter, or kauditd_printk_skb: N callbacks suppressed in the kernel log, means the kernel is discarding records because the backlog is full. Raise -b (the default is 8192) in /etc/audit/rules.d/10-base-config.rules and reload.

Rate limits drop records silently

Nothing on the host reports this failure. Rate limiters sit between auditd and the collector, and they drop messages once execution auditing raises the message rate. Which ones apply depends on how rsyslog reads the records:

grep -rnE 'imuxsock|imjournal|SysSock' /etc/rsyslog.conf /etc/rsyslog.d/
Configuration found The path records take Limiters that apply
imjournal loaded, imuxsock with SysSock.Use="off" — the RHEL family default plugin → journald → rsyslog journald and imjournal
imuxsock reading the socket, no imjournal plugin → rsyslog imuxsock only
Limiter Typical default Effect when exceeded
journald RateLimitIntervalSec=30s, RateLimitBurst=10000 Drops per service, logs a Suppressed N messages line
imjournal Ratelimit.Interval=600, Ratelimit.Burst=20000 Drops before the forwarding action — roughly 33 messages/second sustained
imuxsock Distribution-dependent; check SysSock.RateLimit.Burst Drops before the forwarding action

Each audited process launch produces two or more records, so a host averaging 20 execs per second is already past the imjournal default. The rules still load, auditctl -s still reports lost 0, and the audit log on disk is complete; only the copy Secure60 receives has gaps.

Raising the limits

Both limiters accept 0, meaning unlimited. A finite ceiling is required: the limiter is what stops a runaway logger from filling the disk, saturating the collector, or starving everything else on the host.

The ceiling is set generously — high enough that normal and heavy operation never approach it, low enough that a runaway is still stopped. 300000 messages per 30 seconds is 10,000 a second, roughly twelve times the busiest host in the Secure60 fleet:

# /etc/systemd/journald.conf
RateLimitIntervalSec=30s
RateLimitBurst=300000
# rsyslog — edit the existing module line, do not add a second one
module(load="imjournal" StateFile="imjournal.state"
       Ratelimit.Interval="30" Ratelimit.Burst="300000")

module(load="imuxsock" SysSock.RateLimit.Interval="30" SysSock.RateLimit.Burst="300000")
sudo systemctl restart systemd-journald rsyslog

Both parameter sets are portable across current distributions. RateLimitIntervalSec needs systemd 230 or later, which covers everything from RHEL 8 and Ubuntu 18.04 onward; on anything older the key is RateLimitInterval. The rsyslog parameters have been available since 7.3. The two rsyslog modules name the parameter differently: Ratelimit.Interval for imjournal, SysSock.RateLimit.Interval for imuxsock.

Drop detection

A ceiling is only useful with an alert on it. Every limiter in this path announces itself in the log, and because these hosts forward syslog to Secure60, those announcements arrive as ordinary events a detection can match.

These are the exact strings the shipped binaries emit:

Source Message Means
rsyslogd <module>: N messages lost due to rate-limiting (N allowed within N seconds) imjournal or imuxsock dropped records
rsyslogd <source>: begin to drop messages due to rate-limiting a per-source limiter just started dropping
systemd-journald Suppressed N messages from <unit> journald dropped records for one unit
systemd-journald Forwarding to syslog missed N messages. journald could not hand everything to rsyslog
kernel kauditd_printk_skb: N callbacks suppressed the kernel audit backlog is under pressure

One detection covers all four sources:

message_text : 'messages lost due to rate-limiting'
  OR message_text : 'begin to drop messages due to rate-limiting'
  OR message_text : 'Forwarding to syslog missed'
  OR message_text : 'callbacks suppressed'

On the host itself, the same check:

journalctl --since "-24h" | grep -cE "messages lost due to rate-limiting|Suppressed [0-9]+ messages|Forwarding to syslog missed|callbacks suppressed"
sudo auditctl -s | grep lost

A non-zero count means records were lost for that window.

High-volume hosts: reading the audit log directly

On hosts where execution auditing produces sustained high rates, reading the audit log file directly avoids journald and imjournal altogether:

module(load="imfile")
input(type="imfile"
      File="/var/log/audit/audit.log"
      Tag="audisp-syslog"
      Severity="info"
      Facility="local6")

The syslog plugin is disabled on this route, or every record is sent twice. The Tag stays as shown so events arrive with the app_name the parser template matches on.

A dedicated syslog facility for audit

By default audisp-syslog writes at facility user, priority info, mixed in with everything else on the host. A dedicated facility allows audit to be rate-limited, filtered or routed separately in rsyslog without touching the rest of the log stream:

# /etc/audit/plugins.d/syslog.conf
args = LOG_INFO LOG_LOCAL6

Protecting the ruleset from change

Once the ruleset is settled, -e 2 makes the audit configuration immutable until the host reboots, so auditing cannot be disabled without a reboot that is itself visible. It must be the last line of the merged ruleset, so it goes in its own high-numbered file:

echo '-e 2' | sudo tee /etc/audit/rules.d/99-finalize.rules
sudo augenrules --load

This is applied only after the rules are stable: with -e 2 set, every subsequent rule change needs a reboot.

Existing Auditbeat deployments

Auditbeat remains supported and is documented in the FIM and Audit logs guide. Existing deployments can stay as they are; new Linux hosts use auditd, which needs no third-party agent, uses the syslog path already opened for Linux server logs, and whose rule syntax is the one CIS benchmarks and audit standards are written against.

Back to top