mirror of
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
synced 2026-08-31 02:21:39 -04:00
e566d0625876b699ee0e752d672c4de73a440764
Christian Brauner <brauner@kernel.org> says: binfmt_misc: bpf-backed binary type handlers This is a POC for the nix people and Farid and Eric in particular. I would take my hands off the wheel now that I POCed this and hand it to Farid if he likes to take it forward. VL;MR (very long, must read): For a while now Farid has been trying to make relocatable, hermetic binaries (think Nix-style store layouts) work without patchelf tricks or wrapper scripts. For such binaries the right dynamic loader can only be determined relative to the location of the binary itself, which neither PT_INTERP nor a fixed binfmt_misc interpreter string can express. The first attempt was $ORIGIN expansion in PT_INTERP [1]. I pushed back on that. Userspace guards $ORIGIN behind AT_SECURE so the kernel would have to make the used loader depend on the type of binary, LSMs would need a say, it changes long-standing behavior in ways that are ripe for loader injection attacks, and bprm->file may not have a usable path at all (memfds, deleted files, unresolvable paths). Making the kernel splice bprm->file back together with PT_INTERP is terrible. The second attempt was a pluggable ELF interpreter loader registry [2] which would mean actual kernel modules for custom binary formats. Also no. binfmt_misc was invented to kill exactly this horrendous past. What I suggested instead [3] was to put this where delegating binary formats to userspace already lives: binfmt_misc. The only things binfmt_misc cannot do today are matching programmatically and computing the interpreter per binary instead of using a fixed string recorded at registration time. Farid prototyped that with an eBPF program [4] and it turned out quite workable, but the prototype ran a SOCKET_FILTER program over bprm->buf, added a new helper to the frozen uapi helper list, and returned the computed path through per-CPU memory. This series is the proposal turned into what I think the bpf side {c,sh}ould actually look like. It is a POC: it builds, the selftests pass, and the design is what I want to discuss. The selftests are Farid's from his v2 posting, adapted to the contract below. A handler is an instance of the new binfmt_misc_ops struct_ops with a name and two ops: struct binfmt_misc_ops { bool (*match)(struct linux_binprm *bprm); int (*load)(struct linux_binprm *bprm); char name[BINFMT_MISC_OPS_NAME_MAX]; }; Both programs receive the bprm as a trusted BTF pointer and both are sleepable. The match program decides from the entry lookup walk whether the handler applies, under the same rules as magic matching: registration order, first match wins. It is not limited to the prefetched 256 bytes in bprm->buf: it can read arbitrary file content through bpf_dynptr_from_file(), e.g. to find an ELF interpreter segment at whatever offset it sits. That is what makes multiple independent handlers workable at all - a handler that cannot read the file would have to match broadly and reject from its load program, stealing the binaries of every handler registered after it. To make this safe the entry walk becomes an SRCU read-side section. The load program of the matched handler then selects the interpreter, reading the file the same way and resolving the binary's location via bpf_path_d_path() on &bprm->file->f_path. That also solves the prototype's limitation of only seeing the first 256 bytes of the file. Selecting is the load program's privilege: the verifier rejects the selection kfuncs in match, keyed off the struct_ops member a program attaches to. A match commits the exec to the handler: a failing load fails the exec instead of falling through to later entries, with -ENOEXEC handing over to the remaining binary formats, so the walk is never left and re-entered. The genuinely new piece of bpf surface is a small family of kfuncs: int bpf_binprm_set_interp(struct linux_binprm *bprm, const char *path, size_t path__sz); int bpf_binprm_set_interp_arg(struct linux_binprm *bprm, const char *arg, size_t arg__sz); int bpf_binprm_set_flags(struct linux_binprm *bprm, enum bpf_binprm_flags flags); staging the selected interpreter, an optional single argument for it (the slot the optional argument of a #! interpreter line has), and the per-exec invocation flags - 'P', 'C' and 'O' equivalents. Selection cannot go through bprm_change_interp() directly because load_misc_binary() copies bprm->interp into argv[1] after the program ran, hence the staging fields added in patch 1. Registering (attaching) the struct_ops map publishes the handler under its name in a registry keyed by the registering task's user namespace. Activation reuses the existing text interface with a new 'B' type where the interpreter field carries the handler name - it consistently names whoever supplies the interpreter - and offset, magic, and mask must be empty: echo ':origin:B::::nix:' > /proc/sys/fs/binfmt_misc/register This keeps the existing permission and namespacing model completely intact. Activating a handler requires the same write access to a binfmt_misc instance as any other registration, a container mounting its own instance escapes the host's entries exactly as before, and shadowing e.g. all ELF binaries takes the same privilege as a static 'M' entry matching \x7fELF does today. The only novelty is that matching becomes programmable. Handler lookup walks the user namespace hierarchy upwards, mirroring how binfmt_misc instances themselves are resolved, so a handler registered on the host can be activated from a container's own instance without being forced upon it. The computed interpreter is opened with open_exec() under the caller's credentials and goes through the full LSM vetting as the next binprm level, exactly like a statically registered interpreter, so the program cannot widen access. It only ever redirects the caller to something the caller could exec anyway. A 'B' entry carries no flags in the register string: the load program chooses the invocation flags per exec through bpf_binprm_set_flags() instead. BPF_BINPRM_PRESERVE_ARGV0, BPF_BINPRM_CREDENTIALS and BPF_BINPRM_EXECFD keep the static 'P', 'C' and 'O' semantics - BPF_BINPRM_CREDENTIALS honors the matched binary's suid bits exactly as a static 'C' entry does, with the setuid transition gated by vfsuid_has_mapping() in the caller's user namespace either way, which makes 'B' handlers usable for a per-binary loader over setuid binaries. 'F' (pre-open a fixed interpreter) is rejected: a 'B' entry has no fixed interpreter. AT_EXECVE_CHECK never invokes programs and interpreter chains stay capped by the usual ELOOP depth. A handler for the Nix case then looks roughly like: SEC("struct_ops.s/match") bool BPF_PROG(nix_match, struct linux_binprm *bprm) { return !bpf_strncmp(bprm->buf, 4, "\x7f" "ELF"); } SEC("struct_ops.s/load") int BPF_PROG(nix_load, struct linux_binprm *bprm) { char path[256]; long n; n = bpf_path_d_path(&bprm->file->f_path, path, sizeof(path)); if (n < 0) return n; /* derive the loader location from the binary's path */ return bpf_binprm_set_interp(bprm, path, sizeof(path)); } SEC(".struct_ops.link") struct binfmt_misc_ops nix = { .match = (void *)nix_match, .load = (void *)nix_load, .name = "nix", }; Farid, this should slot underneath your qemu demo from [4] with the program ported to struct_ops. Feel free to take it from here. [1]: https://lore.kernel.org/20260622043934.179879-1-farid.m.zakaria@gmail.com [2]: https://lore.kernel.org/20260702214247.1253741-1-farid.m.zakaria@gmail.com [3]: https://lore.kernel.org/20260703-meditation-ratsuchende-moratorium-9ecdf1f3f8bb@brauner [4]: https://lore.kernel.org/20260704211409.1978485-1-farid.m.zakaria@gmail.com * patches from https://patch.msgid.link/20260714-work-bpf-binfmt_misc-v2-0-57b7529c002c@kernel.org: selftests/exec: add binfmt_misc bpf-backed handler test binfmt_misc: let a bpf handler choose the invocation flags per exec binfmt_misc: let bpf handlers pass an argument to the interpreter bpf: allow fs kfuncs for binfmt_misc_ops programs binfmt_misc: wire up bpf-backed 'B' entries binfmt_misc: let the entry lookup walk sleep binfmt_misc: add binfmt_misc_ops bpf struct_ops exec: stash bpf-selected interpreter state in struct linux_binprm Link: https://patch.msgid.link/20260714-work-bpf-binfmt_misc-v2-0-57b7529c002c@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
Linux kernel ============ The Linux kernel is the core of any Linux operating system. It manages hardware, system resources, and provides the fundamental services for all other software. Quick Start ----------- * Report a bug: See Documentation/admin-guide/reporting-issues.rst * Get the latest kernel: https://kernel.org * Build the kernel: See Documentation/admin-guide/quickly-build-trimmed-linux.rst * Join the community: https://lore.kernel.org/ Essential Documentation ----------------------- All users should be familiar with: * Building requirements: Documentation/process/changes.rst * Code of Conduct: Documentation/process/code-of-conduct.rst * License: See COPYING Documentation can be built with make htmldocs or viewed online at: https://www.kernel.org/doc/html/latest/ Who Are You? ============ Find your role below: * New Kernel Developer - Getting started with kernel development * Academic Researcher - Studying kernel internals and architecture * Security Expert - Hardening and vulnerability analysis * Backport/Maintenance Engineer - Maintaining stable kernels * System Administrator - Configuring and troubleshooting * Maintainer - Leading subsystems and reviewing patches * Hardware Vendor - Writing drivers for new hardware * Distribution Maintainer - Packaging kernels for distros * AI Coding Assistant - LLMs and AI-powered development tools For Specific Users ================== New Kernel Developer -------------------- Welcome! Start your kernel development journey here: * Getting Started: Documentation/process/development-process.rst * Your First Patch: Documentation/process/submitting-patches.rst * Coding Style: Documentation/process/coding-style.rst * Build System: Documentation/kbuild/index.rst * Development Tools: Documentation/dev-tools/index.rst * Kernel Hacking Guide: Documentation/kernel-hacking/hacking.rst * Core APIs: Documentation/core-api/index.rst Academic Researcher ------------------- Explore the kernel's architecture and internals: * Researcher Guidelines: Documentation/process/researcher-guidelines.rst * Memory Management: Documentation/mm/index.rst * Scheduler: Documentation/scheduler/index.rst * Networking Stack: Documentation/networking/index.rst * Filesystems: Documentation/filesystems/index.rst * RCU (Read-Copy Update): Documentation/RCU/index.rst * Locking Primitives: Documentation/locking/index.rst * Power Management: Documentation/power/index.rst Security Expert --------------- Security documentation and hardening guides: * Security Documentation: Documentation/security/index.rst * LSM Development: Documentation/security/lsm-development.rst * Self Protection: Documentation/security/self-protection.rst * Reporting Vulnerabilities: Documentation/process/security-bugs.rst * CVE Procedures: Documentation/process/cve.rst * Embargoed Hardware Issues: Documentation/process/embargoed-hardware-issues.rst * Security Features: Documentation/userspace-api/seccomp_filter.rst Backport/Maintenance Engineer ----------------------------- Maintain and stabilize kernel versions: * Stable Kernel Rules: Documentation/process/stable-kernel-rules.rst * Backporting Guide: Documentation/process/backporting.rst * Applying Patches: Documentation/process/applying-patches.rst * Subsystem Profile: Documentation/maintainer/maintainer-entry-profile.rst * Git for Maintainers: Documentation/maintainer/configure-git.rst System Administrator -------------------- Configure, tune, and troubleshoot Linux systems: * Admin Guide: Documentation/admin-guide/index.rst * Kernel Parameters: Documentation/admin-guide/kernel-parameters.rst * Sysctl Tuning: Documentation/admin-guide/sysctl/index.rst * Tracing/Debugging: Documentation/trace/index.rst * Performance Security: Documentation/admin-guide/perf-security.rst * Hardware Monitoring: Documentation/hwmon/index.rst Maintainer ---------- Lead kernel subsystems and manage contributions: * Maintainer Handbook: Documentation/maintainer/index.rst * Pull Requests: Documentation/maintainer/pull-requests.rst * Managing Patches: Documentation/maintainer/modifying-patches.rst * Rebasing and Merging: Documentation/maintainer/rebasing-and-merging.rst * Development Process: Documentation/process/maintainer-handbooks.rst * Maintainer Entry Profile: Documentation/maintainer/maintainer-entry-profile.rst * Git Configuration: Documentation/maintainer/configure-git.rst Hardware Vendor --------------- Write drivers and support new hardware: * Driver API Guide: Documentation/driver-api/index.rst * Driver Model: Documentation/driver-api/driver-model/driver.rst * Device Drivers: Documentation/driver-api/infrastructure.rst * Bus Types: Documentation/driver-api/driver-model/bus.rst * Device Tree Bindings: Documentation/devicetree/bindings/ * Power Management: Documentation/driver-api/pm/index.rst * DMA API: Documentation/core-api/dma-api.rst Distribution Maintainer ----------------------- Package and distribute the kernel: * Stable Kernel Rules: Documentation/process/stable-kernel-rules.rst * ABI Documentation: Documentation/ABI/README * Kernel Configuration: Documentation/kbuild/kconfig.rst * Module Signing: Documentation/admin-guide/module-signing.rst * Kernel Parameters: Documentation/admin-guide/kernel-parameters.rst * Tainted Kernels: Documentation/admin-guide/tainted-kernels.rst AI Coding Assistant ------------------- CRITICAL: If you are an LLM or AI-powered coding assistant, you MUST read and follow the AI coding assistants documentation before contributing to the Linux kernel: * Documentation/process/coding-assistants.rst This documentation contains essential requirements about licensing, attribution, and the Developer Certificate of Origin that all AI tools must comply with. Communication and Support ========================= * Mailing Lists: https://lore.kernel.org/ * IRC: #kernelnewbies on irc.oftc.net * Bugzilla: https://bugzilla.kernel.org/ * MAINTAINERS file: Lists subsystem maintainers and mailing lists * Email Clients: Documentation/process/email-clients.rst
Description
Languages
C
96.9%
Assembly
0.9%
Rust
0.6%
Shell
0.6%
Python
0.5%
Other
0.3%