Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
30 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -83,3 +83,4 @@ Icon
Network Trash Folder
Temporary Items
.apdisk
.claude/*
84 changes: 84 additions & 0 deletions docs/public/Installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@ This section provides information about the inventory, features, and steps for i
- [CRI](#cri)
- [modprobe](#modprobe)
- [sysctl](#sysctl)
- [fsmount](#fsmount)
- [audit](#audit)
- [Kubernetes Policy](#audit-kubernetes-policy)
- [Daemon](#audit-daemon)
Expand Down Expand Up @@ -2583,6 +2584,89 @@ The following settings are supported in the extended format:

**Warning**: If the changes to the hosts `sysctl` configurations are detected, a reboot is scheduled. After the reboot, the new parameters are validated to match the expected configuration.

#### fsmount

*Installation task*: `prepare.system.fsmount`

*Can cause a reboot*: **Yes** – when `preparation_scripts` contains more than one entry, the node is rebooted between consecutive scripts.

*Can restart a service*: **No**

*Overwrites files*: **Yes** – the rendered systemd unit file is uploaded to the location defined by `template.destination`. A backup of any existing file is retained.

*OS‑specific*: **No**

The `services.fsmount` section allows you to declare additional filesystems that should be formatted and mounted on cluster nodes using **systemd** unit files. By default the section is empty.

Each entry may contain the following keys:

| Parameter | Mandatory | Description |
|------------------------|-----------|-------------|
| **name** | **yes** | Identifier for the mount entry; used in logs and dump filenames. |
| **enabled** | no | Whether the entry is processed. Defaults to **true**. |
| **device** | **yes** | Device file to mount (e.g. `/dev/sdd1`, `/dev/zram0`). |
| **path** | **yes** | Target mount point on the node. |
| **template.source** | **yes** | Path to the Jinja2 template that generates the systemd unit. Can be an internal resource (relative to the Kubemarine package) or an absolute external path. |
| **template.destination**| **yes** | Absolute path on the node where the rendered unit file will be placed. |
| **size** | no | Desired filesystem size (e.g. `1G`). Required for virtual devices such as zram. |
| **type** | no | Filesystem type (e.g. `ext4`, `tmpfs`). |
| **preparation_scripts** | no | Ordered list of shell scripts that run before the systemd unit is installed. Each script may be an internal resource (relative to the Kubemarine package) or an absolute external path. A node reboot is performed between consecutive scripts. If any script exits with a non‑zero status, the procedure fails. |
| **groups** | no | List of node roles (e.g. `control-plane`, `worker`) to which the mount should be applied. |
| **nodes** | no | List of specific node names to which the mount should be applied. |

**Notes**

* You may specify both `groups` and `nodes`; the resulting node set is the union of both selectors.
* If neither `groups` nor `nodes` is provided, the mount is applied to **all** nodes.
* If the mount point already appears in `/proc/mounts`, the entry is skipped for that node.
* Each preparation script is uploaded to the node, executed, and then removed. Scripts are useful for loading kernel modules, upgrading the kernel, or installing prerequisite packages.
* When more than one preparation script is specified, the node is rebooted between consecutive scripts so that kernel or module changes take effect before the next script runs.
* If `enabled` is set to **false**, the mount entry is ignored entirely.

**Warning**: Failure of any `preparation_scripts` entry is fatal — the whole procedure stops immediately.

The following variables are made available to the Jinja2 template:

| Variable | Source |
|----------|--------|
| `name` | `name` field |
| `device` | `device` field |
| `path` | `path` field |
| `size` | `size` field (empty string if omitted) |
| `type` | `type` field (empty string if omitted) |

This functionality could be used to mount ZRAM volume on all cluster nodes for pods logs. The cluster.yaml part is the following:

```yaml
services:
fsmount:
- name: zram
size: 1G
type: ext4
device: /dev/zram0
path: /var/log/pods
template:
source: templates/zram-setup.service.j2
destination: /etc/systemd/system/zram-setup.service
preparation_scripts:
- resources/scripts/upgrade_kernel.sh
- resources/scripts/zram.sh
groups: [control-plane, worker]
```
Comment thread
theboringstuff marked this conversation as resolved.

Comment thread
theboringstuff marked this conversation as resolved.
The node is rebooted between the two scripts so the upgraded kernel is running before `zram.sh` executes.

The `size` of the ZRAM volume must be selected according to the kubelet configuration (`containerLogMax` options). For the current case, the following values are recommended and they must be set in the `kubeadm_kubelet` section:

```yaml
services:
kubeadm_kubelet:
containerLogMaxSize: 5Mi
containerLogMaxFiles: 2
```

**Warning**: Pay attention, the OS must provide `zram` module.

#### audit

##### Audit Kubernetes Policy
Expand Down
17 changes: 17 additions & 0 deletions docs/public/Kubecheck.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,7 @@ This section provides information about the Kubecheck functionality.
- [215 Firewalld Status](#215-firewalld-status)
- [216 Swap State](#216-swap-state)
- [217 Modprobe Rules](#217-modprobe-rules)
- [236 Filesystem Mounts](#236-filesystem-mounts)
- [218 Time Difference](#218-time-difference)
- [219 Health Status ETCD](#219-health-status-etcd)
- [220 Control Plane Configuration Status](#220-control-plane-configuration-status)
Expand Down Expand Up @@ -657,6 +658,22 @@ The test verifies that swap is disabled on all nodes in the cluster, otherwise t
The test compares the modprobe rules on the nodes with the rules specified in the inventory or with default rules. If
rules does not match, the test will fail.

##### 236 Filesystem Mounts

*Task*: `services.system.fsmount.mounts`

This check validates that every **enabled** filesystem mount defined in ``services.fsmount`` is correctly configured on ``control‑plane`` and ``worker`` nodes. Entries with ``enabled: false`` are ignored.

For each applicable entry the following conditions are verified:

* The mount point specified by ``path`` appears in ``/proc/mounts``.
* If the ``type`` field is defined in the inventory, the filesystem type reported in ``/proc/mounts`` must match it. When ``type`` is omitted only the presence of the mount point is checked.
* For ZRAM‑backed devices (i.e., the ``device`` value starts with ``/dev/zram``), the mount point must also be listed in the output of ``zramctl --output‑all``.

If any validation fails, the test reports the affected node and mount path with a detailed error message.

**Note**: Nodes that have no applicable ``services.fsmount`` entries are skipped silently.

##### 218 Time Difference

*Task*: `services.system.time`
Expand Down
85 changes: 85 additions & 0 deletions docs/public/Maintenance.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@ This section describes the features and steps for performing maintenance procedu
- [Reconfigure Procedure](#reconfigure-procedure)
- [Manage PSS Procedure](#manage-pss-procedure)
- [Reboot Procedure](#reboot-procedure)
- [Mount Filesystems Procedure](#mount-filesystems-procedure)
- [Certificate Renew Procedure](#certificate-renew-procedure)
- [Etcd Member Reunion](#etcd-member-reunion)
- [Procedure Execution](#procedure-execution)
Expand Down Expand Up @@ -1357,6 +1358,90 @@ nodes:
```


## Mount Filesystems Procedure

The `mount_fs` procedure configures filesystems on cluster nodes based on the supplied `procedure.yaml`.

For each node that contains at least one applicable **fsmount** entry, the behaviour depends on the value of the
``reboot`` option:

**When ``reboot: true`` (the default):**
1. Drains the node (if it is a ``control-plane`` or ``worker``) to safely evacuate workloads.
2. Removes any existing data from each configured mount path, ensuring a clean state.
3. Installs and enables the corresponding systemd mount units.
4. Reboots the node so the new mounts become active at the operating‑system level.
5. Uncordons the node (again, for ``control-plane`` or ``worker``) to make it schedulable.

**When ``reboot: false``:**
1. Removes existing data from each configured mount path.
2. Installs and enables the systemd mount units immediately, without performing a reboot.

Nodes that have no applicable items are skipped entirely. After a successful execution, ``cluster.yaml`` is updated
with the fsmount items defined in the procedure file.

**Note**: Data inside the configured mount paths is erased before the mount is set up. Back up any important data
prior to running this procedure.
**Note**: For the case when ZRAM volume is mounted to `/var/log/pods`, the kubelet service must be reconfigured
with the recomended parameters. The `procedure.yaml` is as follows:

```yaml
services:
kubeadm_kubelet:
containerLogMaxSize: 5Mi
containerLogMaxFiles: 2
```

**Warning**: Pay attention, the OS must provide `zram` module.

### Mount Filesystems Procedure Parameters

The procedure requires a positional argument that points to the procedure inventory file.

The JSON schema for the inventory is available at
[URL](/kubemarine/resources/schemas/mount_fs.json?raw=1). See
[Validation by JSON Schemas](Installation.md#inventory-validation) for further details.

#### ``reboot`` Parameter

Controls whether each node is drained and rebooted after the mount units are installed.

* ``true`` (default) – drain, install, reboot, then uncordon. Use this when the filesystem must be cleanly initialised before any
workloads run on the node.
* ``false`` – install and enable the units in‑place without a reboot. Use this when the filesystem can be activated live.

#### ``fsmount`` Parameter

The list of filesystem mount items to apply. Its structure is identical to the ``services.fsmount`` section in ``cluster.yaml``.
For a description of the available fields, refer to the [fsmount](Installation.md#fsmount) section in the _Kubemarine Installation Procedure_.

**Example:**

```yaml
reboot: true
fsmount:
- name: zram
size: 1G
type: ext4
device: /dev/zram0
path: /var/log/pods
template:
source: templates/zram-setup.service.j2
destination: /etc/systemd/system/zram-setup.service
preparation_scripts:
- resources/scripts/upgrade_kernel.sh
- resources/scripts/zram.sh
groups: [control-plane, worker]
```

Comment thread
theboringstuff marked this conversation as resolved.
The node is rebooted between `upgrade_kernel.sh` and `zram.sh` so the new kernel is active before the zram module is loaded.

### Mount Filesystems Procedure Tasks Tree

The ``mount_fs`` procedure executes the following sequence of tasks:

* ``mount_filesystems``
* ``overview``

## Certificate Renew Procedure

The `cert_renew` procedure allows you to renew some certificates on an existing Kubernetes cluster.
Expand Down
8 changes: 6 additions & 2 deletions kubemarine/__main__.py
Original file line number Diff line number Diff line change
Expand Up @@ -50,8 +50,8 @@
if (1, 1) in ir[0] and 'tag:yaml.org,2002:float' in ir[1]:
float_patched_resolver = (ir[1], ir[2], ir[3])
# Globally change behaviour of yaml.safe_load and yaml.dump
yaml.Dumper.add_implicit_resolver(*float_patched_resolver) # type: ignore[no-untyped-call]
yaml.SafeLoader.add_implicit_resolver(*float_patched_resolver) # type: ignore[no-untyped-call]
yaml.Dumper.add_implicit_resolver(*float_patched_resolver) # type: ignore[no-untyped-call, unused-ignore]
yaml.SafeLoader.add_implicit_resolver(*float_patched_resolver) # type: ignore[no-untyped-call, unused-ignore]
break


Expand Down Expand Up @@ -97,6 +97,10 @@
'description': "Renew certificates on Kubernetes cluster",
'group': 'maintenance'
},
'mount_fs': {
'description': "Mount filesystems described in the fsmount inventory section",
'group': 'maintenance'
},
'reboot': {
'description': "Reboot Kubernetes nodes",
'group': 'maintenance'
Expand Down
3 changes: 3 additions & 0 deletions kubemarine/core/resources.py
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@
import kubemarine.keepalived
import kubemarine.kubernetes
import kubemarine.kubernetes_accounts
import kubemarine.fsmount
import kubemarine.modprobe
import kubemarine.packages
import kubemarine.plugins
Expand Down Expand Up @@ -408,6 +409,7 @@ def enrichment_functions(self) -> List[c.EnrichmentFunction]:
kubemarine.cri.enrich_upgrade_inventory,
kubemarine.plugins.nginx_ingress.cert_renew_enrichment,
kubemarine.plugins.envoy_gateway.cert_renew_enrichment,
kubemarine.fsmount.enrich_procedure_inventory,
kubemarine.sysctl.enrich_reconfigure_inventory,
kubemarine.core.inventory.enrich_reconfigure_inventory,
# Enrichment of procedure inventory should be finished at this step.
Expand Down Expand Up @@ -485,6 +487,7 @@ def enrichment_functions(self) -> List[c.EnrichmentFunction]:
kubemarine.system.verify_inventory,
kubemarine.system.enrich_etc_hosts,
kubemarine.modprobe.enrich_kernel_modules,
kubemarine.fsmount.enrich_inventory,

# Calculate some differences between previous and new inventory
# Depends on kubemarine.packages.enrich_inventory
Expand Down
Loading
Loading