fscrypt v1 policies bind master keys to the calling process's keyring,
which means files in a homed-managed directory aren't readable when
accessed through a container bind mount, a different mount namespace,
or by any process other than the one that first unlocked the home.
Reading a file from outside such a context first warms the buffer cache
and papers over the symptom (#18280), but the underlying problem (the
key not flowing across keyrings) remains.
v2 policies (Linux 5.4+) route master keys through the filesystem
keyring via FS_IOC_ADD_ENCRYPTION_KEY / FS_IOC_REMOVE_ENCRYPTION_KEY,
so the key is visible to every process accessing the filesystem.
Switch homed to v2:
- Read the existing policy via FS_IOC_GET_ENCRYPTION_POLICY_EX, which
reports both v1 and v2 policies. The ioctl is available since Linux
5.4, i.e. on every kernel we support (our baseline is 5.10), so no
fallback to the v1-only FS_IOC_GET_ENCRYPTION_POLICY is needed.
- Drive slot matching off a full fscrypt_key_specifier (HomeSetup now
carries that instead of a bare 8-byte descriptor). Slot decryption
derives either the v1 descriptor (SHA-512 double hash) or the v2
identifier (HKDF-SHA512 with the kernel's info string), and compares
against the policy.
- Install the master key the right way per version: add_key("logon", ...)
to thread+user keyrings for v1, FS_IOC_ADD_ENCRYPTION_KEY for v2.
- home_flush_keyring_fscrypt opens the image directory, looks up the
policy, and either calls FS_IOC_REMOVE_ENCRYPTION_KEY (v2) or walks
the user keyring (v1).
- New homes default to v2; fall back to v1 only if the kernel rejects
FS_IOC_ADD_ENCRYPTION_KEY with ENOTTY/EOPNOTSUPP. The v2 create path
derives the identifier locally first, passes it to ADD_KEY as
expected_identifier, and cleans up via REMOVE_KEY if SET_POLICY then
fails, so the v1 fallback never sees a stranded key. Existing v1
homes continue to unlock, rekey, and deactivate as before.
- A v2 master key installed to work on an inactive home is always
removed again unless that home ends up activated. v2 keys persist in
the filesystem keyring until removed explicitly (v1 keys instead died
with the homework process' keyring), so a key left behind would leave
a home nobody activated readable until the next deactivation or reboot.
home_setup_fscrypt() and home_create_fscrypt() therefore arm a rollback
right after installing the key; the activation path disarms it once the
mount is in place (the live home owns the key), while every other path
-- create, and passwd/update/resize of an inactive home, plus all error
paths -- rolls it back via home_setup_done(). Activation reinstalls the
key.
The v1 and v2 on-disk policy formats differ (v1: 8-byte descriptor;
v2: 16-byte identifier) and fscrypt has no in-place upgrade path, so
a v1 home is always unlocked via the v1 code path and a v2 home is
always unlocked via the v2 code path, regardless of kernel version.
Slot xattr format is unchanged: the master key is the same, only how it
binds to the directory changes.
Fixes#18280.
Ananth Bhaskararaman antsub@gmail.com
The new SD_ELF_NOTE_DLOPEN_ANCHORED() macro utilizes the 'o' assembler
flag (SHF_LINK_ORDER), which requires binutils >= 2.35 or LLVM >= 18.
If an LLVM version older than 18 is encountered, the macro automatically
falls back to the non-anchored variant.
This non-anchored fallback relies on the 'R' (SHF_GNU_RETAIN) flag,
which requires binutils >= 2.36 or LLVM >= 13.
Since the codebase now unconditionally adopts SD_ELF_NOTE_DLOPEN_ANCHORED()
tree-wide, the effective minimum toolchain requirements become:
- binutils >= 2.35 (to support the 'o' flag)
- LLVM/Clang >= 13 (to support the 'R' flag fallback for versions < 18)
Update the minimal toolchain versions in the README to reflect these
requirements for building systemd.
The idea is that we can build a container by building a single-binary
systemd:
```console
$ meson setup build-static --default-library=static --prefer-static --auto-features=disabled -Dbuild-static=true -Dsystemd-multicall-binary=true && ninja -C build-static systemd
$ mkdir /var/tmp/container/usr/lib -p
$ cp build-static/systemd /var/tmp/container/usr/lib/
$ echo 'ID=quick' >/var/tmp/container/usr/lib/os-release
$ systemd-nspawn --restrict-address-families=af_unix --register=no --private-users=managed -D /var/tmp/container/ /usr/lib/systemd
░ Spawning container container on /var/tmp/container.
░ Press Ctrl-] three times within 1s to kill container; two times followed by r
░ to reboot container; two times followed by p to poweroff container.
Selected user namespace base 1855193088 and range 65536.
systemd 262~devel running in system mode (-PAM -AUDIT +SELINUX -APPARMOR +IMA +IPE +SMACK -SECCOMP -GCRYPT +GNUTLS +OPENSSL -ACL +BLKID +CURL -ELFUTILS -FIDO2 +IDN2 +KMOD +LIBCRYPTSETUP +LIBCRYPTSETUP_PLUGINS +LIBFDISK +PCRE2 -PWQUALITY +P11KIT +QRENCODE +TPM2 -BZIP2 -LZ4 +XZ +ZLIB +ZSTD -BPF_FRAMEWORK -BTF -XKBCOMMON +UTMP -LIBARCHIVE)
Detected virtualization systemd-nspawn.
Detected architecture x86-64.
Detected first boot.
Welcome to Linux!
Initializing machine ID from container UUID.
Failed to open netlink, ignoring: Address family not supported by protocol
Applying preset policy.
Populated /etc with preset unit settings.
Unit default.target not found.
Falling back to graphical.target.
Mount unit not supported, skipping *MountsFor= dependencies.
Queued start job for default target graphical.target.
[ OK ] Reached target sysinit.target.
[ OK ] Reached target basic.target.
System is tainted: unmerged-bin:var-run-bad
[ OK ] Reached target multi-user.target.
[ OK ] Reached target graphical.target.
Startup finished in 61ms.
```
The container can be reloaded with SIGTERM, powered off with SIGRTMIN+4,
etc. SIGRTMIN+5 should cause a reboot but it currently fails:
```
...
Rebooting.
Container container is being rebooted.
Failed to attach root directory: Invalid argument
Failed to receive mount namespace fd from outer child: Input/output error
```
It's a bug … somewhere, but probably not caused by the linking changes
being done here.
With gpg sub keys one can rotate signing keys while having a stable
trust anchor. So far one still had to ship the sub key out of band but
a newer gpg has the option to include the sub key in the signature and
import it automatically. This is safe if we only allow importing a sub
key signed by the top key we already have in the key ring.
Add the --auto-key-import argument to gpg to import subkeys but also
set --import-options=merge-only,import-clean to restrict what we import
to only be sub keys signed by the top key we have in the keyring and
discard any irrelevant parts. The ugly part is that we also have to
work on a temporary copy of the keyring because gpg wants to persist
the added key material but we don't what that here.
Add support for two newer NUMA memory policies:
- MPOL_PREFERRED_MANY (Linux 5.15): like MPOL_PREFERRED but accepts
a set of nodes instead of a single node, falling back to all nodes
if preferred nodes cannot satisfy the allocation.
- MPOL_WEIGHTED_INTERLEAVE (Linux 6.9): like MPOL_INTERLEAVE but
distributes pages across nodes proportionally to per-node weights
configured via /sys/kernel/mm/mempolicy/weighted_interleave/.
On kernels that do not support the requested policy, set_mempolicy()
returns EINVAL. We convert EINVAL to EOPNOTSUPP only for the two new
policies (MPOL_PREFERRED_MANY, MPOL_WEIGHTED_INTERLEAVE), so that a
bad NUMAMask= for already-supported policies still fails the service
rather than being silently ignored.
The NUMA subsystem being absent (ENOSYS) continues to be handled
silently at debug level, as before.
Varlink serialization uses json_underscorify() on an owned copy of
the policy name string to convert hyphenated names to the underscore
form declared in the IDL enum, avoiding mutation of the read-only
static string table.
Signed-off-by: dongshengyuan <dongshengyuan@uniontech.com>
With this change, gcrypt dependency is not mandatory. Hence, allow to build
systemd even when -D gcrypt=enabled but gcrypt devel package is not installed.
Because fdisk_assign_device tries to open block devices with O_EXCL, when it
does it blocks cryptsetup from using partition block devices for the same
disk.
Since we already have a file descriptor for the device, we can just share it
and use fdisk_assign_device_by_fd instead.
This requires at least libfdisk 2.35 (part of util-linux) which was
released in 2020.
This baseline bump is mainly to support the secure mode feature
in more(1) that has been made available since util-linux v2.42.
Signed-off-by: Christian Goeschel Ndjomouo <cgoesc2@wgu.edu>
It was bumped in a40d934007 but this
is hardly load bearing stuff so let's document the version we actually
require rather than the version that makes a hardly load bearing feature
work properly, especially since v2.41 is extremely new and requiring
distributions to have that is just unrealistic.
This doesn't actually change anything materially except documentation,
but it keeps us honest about depending on stuff from newer util-linux
because we happen to document reliance on an extremely new version.
agetty from util-linux is meanwhile following the configuration file
specification for /etc/issue. The usage of "--issue-file" breaks this
on distributions with current util-linux.
The previous minimum required version 5.4 will be EOL on 2025-12.
Let's bump the required minimum kernel version to the next LTS release 5.10
(released on 2020-12-13, EOL on 2026-12, CIP support until 2031-01).
The new recommended baseline 5.14 is the version that CentOS 9 uses.
CentOS 9 will EOL on 2027-05.
See also #38608.
The current tree doesn't even compile with libidn(1) after
2c7bdaf9f1, which included
a non-existent call to check_dlopen_blocked() somehow.
Hence, it feels safe to just nuke legacy support from
our repo.
Note, this drops logging only test case for crypt_preferred_method(),
as that requires explicitly dlopen() the library. But, we should test
that make_salt() and friends automatically dlopen() it.
libcrypt was no longer built by default since glibc-2.38, and it has been
completely removed since glibc-2.39.
Let's always use libxcrypt, unless when building with musl. As already
major distribution already have libxcrypt-4.4.x, hence let's also bump
the required minimum version to 4.4.0.
libxcrypt cannot be built with musl, hence the previous fallback logic
in libcrypt-util.c are moved to musl/crypt.c.
Note, libxcrypt-4.4.0 was released on 2018-11-20.
See also #38608.
Major distributions already have libseccomp 2.5.x or newer.
Let's bump to the required minimum version to 2.4.0, which provides
SCMP_ACT_KILL_PROCESS, SCMP_ACT_LOG, SCMP_ARCH_PARISC, and
SCMP_ARCH_PARISC64.
Note, libseccomp 2.4.0 was released on 2019-03-15.
See also #38608.
Major distributions already have cryptsetup newer than 2.4.0.
Let's bump the minimal required version.
Note, cryptsetup 2.4.0 was released on 2021-08-18.
See also #38608.
Major distributions already have elfutils >= 0.190.
Let's bump the required minimum version.
Note, elfutils 0.177 was released on 2019-08-14.
See also #38608.
Major distributions already have blkid >= 2.37.
Let's bump the minimal required version.
Note, util-linux (which provides blkid) 2.37 was released on 2021-06-01.
See also #38608.
All major distributions have switched to OpenSSL version 3.x.
Let's drop support of OpenSSL version 1.x.
Note, OpenSSL 3.0 was released on 2021-09-07 (and will be EOL on 2026-09-07).
See also #38608.
meson test output is extremely verbose, printing
a separate line for each successful test. Let's
add -q/--quiet everywhere so it only prints full
lines for skipped and failed tests.
It helps nobody to break compatibility for a missing definition
for printing an error.
Just add the missing definition if not present, as it is already
done for thousands of others from the kernel, glibc, etc.
This partially reverts commit d8b60944f5.
Major distributions already have libfido2 >= 1.12.0.
Let's bump the required minimum version to 1.5.0, which provides
FIDO_ERR_UV_BLOCKED.
Note, libfido2 1.5.0 was released on 2020-09-01.
See also #38608.
Most compat glue has been already removed, except for several cgroup v1
specific codes. It is too late to remove the remaining things before v258.
Let's remove them after v258.
We require at least crypt_r() exists, and it is provided since glibc-2.0
(and dropped in glibc-2.39) or by libxcrypt, and the function is
provided in crypt.h regardless it is provided by glibc or libxcrypt.
Hence, we cannot fallback to unistd.h.
This makes the condition about crypt.h more strict, and stop compilation
earlier when crypt.h does not exist.