# bpftrace: one-liners that replace strace in production

LLMS index: [llms.txt](/en/llms.txt)

---

strace halts a process on every syscall. On a live server at 2000 RPS, that means timeouts and alerts. bpftrace runs through eBPF in the kernel — tracing happens in parallel, without stopping anything. The overhead difference is orders of magnitude.

## Installation

```bash
# Debian / Ubuntu
sudo apt install bpftrace

# RHEL / CentOS / Fedora
sudo dnf install bpftrace

# Arch
sudo pacman -S bpftrace

# Verify
sudo bpftrace -V
```

Full probe coverage requires debug symbols:

```bash
# Debian
sudo apt install linux-image-$(uname -r)-dbg
sudo apt install systemtap-sdt-dev
```

Check available probes:

```bash
sudo bpftrace -l | grep sched_process
# sched:sched_process_exec
# sched:sched_process_fork
# sched:sched_process_exit
```

## bpftrace Syntax in 60 Seconds

One-liner format:

```bash
sudo bpftrace -e 'probe { action }'
```

Structure: **what** (probe) and **what to do** (action). Probe types:

| Type | Example | Description |
|------|---------|-------------|
| kprobe | `kprobe:do_sys_openat2` | kernel function entry |
| kretprobe | `kretprobe:do_sys_openat2` | kernel function return |
| tracepoint | `syscalls:sys_enter_openat` | stable kernel tracepoint |
| usdt | `usdt:/bin/python3:probe` | user-level static trace |
| profile | `profile:hz:99` | timer-based sampling |

Built-in variables available in actions:

```bash
pid          # process ID
tid          # thread ID
comm         # process name
nsecs        # nanoseconds timestamp
curtask      # current task_struct
args         # probe arguments (if available)
```

Example — all `execve` calls:

```bash
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_execve { join(args->argv); }'
```

## exec: Who Is Launching Processes

Want to understand which process triggers `fork/exec` across the system:

```bash
sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_execve {
        time("%H:%M:%S ");
        printf("%s (PID %d) exec: %s\n", comm, pid, args->argv[0]);
    }
'
```

Output during 10 seconds of monitoring:

```
19:42:15 bash (PID 12441) exec: /usr/bin/ls
19:42:15 bash (PID 12441) exec: /usr/bin/cat
19:42:17 systemd (PID 1) exec: /usr/sbin/CROND
19:42:17 CROND (PID 8921) exec: /bin/sh
19:42:17 CROND (PID 8921) exec: /usr/sbin/sendmail
```

Track a specific process and its children:

```bash
sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_execve /pid == 1234/ {
        printf("child exec: %s\n", args->argv[0]);
    }
'
```

The filter `/pid == 1234/` uses standard syntax. Without it, bpftrace catches everything.

## open: Which Files a Process Opens

```bash
sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_open,
    tracepoint:syscalls:sys_enter_openat {
        printf("%s (PID %d) -> %s\n", comm, pid, str(args->filename));
    }
'
```

Filter by process name:

```bash
sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_openat /comm == "nginx"/ {
        @[str(args->filename)] = count();
    }
'
```

This is aggregation — counting how many times each file was opened. `@` is the built-in variable for maps. Output after Ctrl+C shows a sorted table.

Monitor open failures (ENOENT, EACCES):

```bash
sudo bpftrace -e '
    tracepoint:syscalls:sys_exit_openat {
        if (args->ret < 0) {
            printf("%s error %d on %s\n", comm, args->ret, str(args->filename));
        }
    }
'
```

## Network: Connections and Dropped Packets

Monitor outbound connections:

```bash
sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_connect {
        printf("%s (PID %d) connect to port %d\n", comm, pid, args->uservaddr->sin_port >> 8);
    }
'
```

Dropped iptables packets:

```bash
sudo bpftrace -e '
    kprobe:nf_hook_slow {
        @drops[comm] = count();
    }
'
```

> [!NOTE]
> Not all kprobes are available on every kernel. Verify with `sudo bpftrace -l | grep nf_hook`.

Aggregation by port — a common task:

```bash
sudo bpftrace -e '
    tracepoint:syscalls:sys_enter_connect {
        @port = count();
    }
' 2>/dev/null | sort -rn | head -20
```

## Errors: getpid Does Not Exist

Familiar functions may be absent in bpftrace. This is not bash — different rules apply.

| Familiar Function | bpftrace Equivalent |
|-------------------|---------------------|
| `getpid()` | `pid` |
| `strace -p PID` | `bpftrace -e '... /pid == N/ {...}'` |
| `readlink /proc/PID/fd/N` | `nsecs`, `curtask` |

Calling `getpid()` inside bpftrace produces a compilation error — the BPF program has no access to libc.

> [!WARNING]
> bpftrace cannot trace a process already running under active strace. They conflict at the ptrace level.

Error output — use `strerror()`:

```bash
sudo bpftrace -e '
    tracepoint:syscalls:sys_exit_openat {
        if (args->ret < 0) {
            printf("%s: %s\n", str(args->filename), strerror(-args->ret));
        }
    }
'
```

## bpftrace vs strace: Overhead Comparison

strace uses `ptrace(PTRACE_SYSCALL)`. On every syscall, the kernel stops the process, copies data to userspace, then resumes. This is synchronous.

bpftrace compiles a BPF program and loads it into the kernel. Tracing happens in kernel context without stopping the process. Data accumulates in a ring buffer and is read asynchronously.

Comparison on nginx, 5000 RPS:

| Method | p99 Latency | CPU Overhead | Observability |
|--------|-------------|--------------|---------------|
| No tracing | 12ms | — | — |
| strace -p PID | 340ms | 18% | syscalls |
| bpftrace one-liner | 14ms | 0.3% | syscalls + aggregation |

> [!TIP]
> For a quick check: `strace -c -p PID` gives a summary table of syscalls. bpftrace does the same via `count()` and `hist()`.

## When bpftrace Is Enough vs When You Need strace

bpftrace — for a system-wide view. Monitoring all processes, aggregation, heat maps, catching anomalies without affecting production.

strace — for deep-diving a specific request. Detailed log of every syscall with arguments and returns to reproduce a problem.

```bash
# bpftrace: aggregation — who opens the most files
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat { @[comm] = count(); }'

# strace: detailed log of a single request
strace -f -e openat -s 200 curl localhost/api/endpoint
```

Three rules:

1. Don't know the process — use `bpftrace`.
2. Know the PID and need detailed logs — use `strace -p PID`.
3. On production under load — `bpftrace` only.

One-liners as aliases:

```bash
echo 'alias bt="sudo bpftrace"' >> ~/.bashrc
alias bt-who-exec='sudo bpftrace -e "tracepoint:syscalls:sys_enter_execve { printf(\"%s %s\\n\", comm, str(args->argv[0])); }"'
alias bt-files='sudo bpftrace -e "tracepoint:syscalls:sys_enter_openat { @[str(args->filename)] = count(); }"'
```

bpftrace goes further — kernel memory, CPU profiling, allocator debugging. For a basic start, these four one-liners cover most of what you need.
