diff --git a/.gitignore b/.gitignore index 5584798f9..45c8125eb 100644 --- a/.gitignore +++ b/.gitignore @@ -83,3 +83,4 @@ Icon Network Trash Folder Temporary Items .apdisk +.claude/* diff --git a/docs/public/Installation.md b/docs/public/Installation.md index 93c3e1b50..348b082a5 100644 --- a/docs/public/Installation.md +++ b/docs/public/Installation.md @@ -46,6 +46,7 @@ This section provides information about the inventory, features, and steps for i - [CRI](#cri) - [modprobe](#modprobe) - [sysctl](#sysctl) + - [fsmount](#fsmount) - [audit](#audit) - [Kubernetes Policy](#audit-kubernetes-policy) - [Daemon](#audit-daemon) @@ -2583,6 +2584,89 @@ The following settings are supported in the extended format: **Warning**: If the changes to the hosts `sysctl` configurations are detected, a reboot is scheduled. After the reboot, the new parameters are validated to match the expected configuration. +#### fsmount + +*Installation task*: `prepare.system.fsmount` + +*Can cause a reboot*: **Yes** – when `preparation_scripts` contains more than one entry, the node is rebooted between consecutive scripts. + +*Can restart a service*: **No** + +*Overwrites files*: **Yes** – the rendered systemd unit file is uploaded to the location defined by `template.destination`. A backup of any existing file is retained. + +*OS‑specific*: **No** + +The `services.fsmount` section allows you to declare additional filesystems that should be formatted and mounted on cluster nodes using **systemd** unit files. By default the section is empty. + +Each entry may contain the following keys: + +| Parameter | Mandatory | Description | +|------------------------|-----------|-------------| +| **name** | **yes** | Identifier for the mount entry; used in logs and dump filenames. | +| **enabled** | no | Whether the entry is processed. Defaults to **true**. | +| **device** | **yes** | Device file to mount (e.g. `/dev/sdd1`, `/dev/zram0`). | +| **path** | **yes** | Target mount point on the node. | +| **template.source** | **yes** | Path to the Jinja2 template that generates the systemd unit. Can be an internal resource (relative to the Kubemarine package) or an absolute external path. | +| **template.destination**| **yes** | Absolute path on the node where the rendered unit file will be placed. | +| **size** | no | Desired filesystem size (e.g. `1G`). Required for virtual devices such as zram. | +| **type** | no | Filesystem type (e.g. `ext4`, `tmpfs`). | +| **preparation_scripts** | no | Ordered list of shell scripts that run before the systemd unit is installed. Each script may be an internal resource (relative to the Kubemarine package) or an absolute external path. A node reboot is performed between consecutive scripts. If any script exits with a non‑zero status, the procedure fails. | +| **groups** | no | List of node roles (e.g. `control-plane`, `worker`) to which the mount should be applied. | +| **nodes** | no | List of specific node names to which the mount should be applied. | + +**Notes** + +* You may specify both `groups` and `nodes`; the resulting node set is the union of both selectors. +* If neither `groups` nor `nodes` is provided, the mount is applied to **all** nodes. +* If the mount point already appears in `/proc/mounts`, the entry is skipped for that node. + * Each preparation script is uploaded to the node, executed, and then removed. Scripts are useful for loading kernel modules, upgrading the kernel, or installing prerequisite packages. + * When more than one preparation script is specified, the node is rebooted between consecutive scripts so that kernel or module changes take effect before the next script runs. + * If `enabled` is set to **false**, the mount entry is ignored entirely. + +**Warning**: Failure of any `preparation_scripts` entry is fatal — the whole procedure stops immediately. + +The following variables are made available to the Jinja2 template: + +| Variable | Source | +|----------|--------| +| `name` | `name` field | +| `device` | `device` field | +| `path` | `path` field | +| `size` | `size` field (empty string if omitted) | +| `type` | `type` field (empty string if omitted) | + +This functionality could be used to mount ZRAM volume on all cluster nodes for pods logs. The cluster.yaml part is the following: + +```yaml +services: + fsmount: + - name: zram + size: 1G + type: ext4 + device: /dev/zram0 + path: /var/log/pods + template: + source: templates/zram-setup.service.j2 + destination: /etc/systemd/system/zram-setup.service + preparation_scripts: + - resources/scripts/upgrade_kernel.sh + - resources/scripts/zram.sh + groups: [control-plane, worker] +``` + +The node is rebooted between the two scripts so the upgraded kernel is running before `zram.sh` executes. + +The `size` of the ZRAM volume must be selected according to the kubelet configuration (`containerLogMax` options). For the current case, the following values are recommended and they must be set in the `kubeadm_kubelet` section: + +```yaml +services: + kubeadm_kubelet: + containerLogMaxSize: 5Mi + containerLogMaxFiles: 2 +``` + +**Warning**: Pay attention, the OS must provide `zram` module. + #### audit ##### Audit Kubernetes Policy diff --git a/docs/public/Kubecheck.md b/docs/public/Kubecheck.md index 41637950b..e2fbb0147 100644 --- a/docs/public/Kubecheck.md +++ b/docs/public/Kubecheck.md @@ -67,6 +67,7 @@ This section provides information about the Kubecheck functionality. - [215 Firewalld Status](#215-firewalld-status) - [216 Swap State](#216-swap-state) - [217 Modprobe Rules](#217-modprobe-rules) + - [236 Filesystem Mounts](#236-filesystem-mounts) - [218 Time Difference](#218-time-difference) - [219 Health Status ETCD](#219-health-status-etcd) - [220 Control Plane Configuration Status](#220-control-plane-configuration-status) @@ -657,6 +658,22 @@ The test verifies that swap is disabled on all nodes in the cluster, otherwise t The test compares the modprobe rules on the nodes with the rules specified in the inventory or with default rules. If rules does not match, the test will fail. +##### 236 Filesystem Mounts + +*Task*: `services.system.fsmount.mounts` + +This check validates that every **enabled** filesystem mount defined in ``services.fsmount`` is correctly configured on ``control‑plane`` and ``worker`` nodes. Entries with ``enabled: false`` are ignored. + +For each applicable entry the following conditions are verified: + +* The mount point specified by ``path`` appears in ``/proc/mounts``. +* If the ``type`` field is defined in the inventory, the filesystem type reported in ``/proc/mounts`` must match it. When ``type`` is omitted only the presence of the mount point is checked. +* For ZRAM‑backed devices (i.e., the ``device`` value starts with ``/dev/zram``), the mount point must also be listed in the output of ``zramctl --output‑all``. + +If any validation fails, the test reports the affected node and mount path with a detailed error message. + +**Note**: Nodes that have no applicable ``services.fsmount`` entries are skipped silently. + ##### 218 Time Difference *Task*: `services.system.time` diff --git a/docs/public/Maintenance.md b/docs/public/Maintenance.md index 892c48bc8..18b334c27 100644 --- a/docs/public/Maintenance.md +++ b/docs/public/Maintenance.md @@ -14,6 +14,7 @@ This section describes the features and steps for performing maintenance procedu - [Reconfigure Procedure](#reconfigure-procedure) - [Manage PSS Procedure](#manage-pss-procedure) - [Reboot Procedure](#reboot-procedure) + - [Mount Filesystems Procedure](#mount-filesystems-procedure) - [Certificate Renew Procedure](#certificate-renew-procedure) - [Etcd Member Reunion](#etcd-member-reunion) - [Procedure Execution](#procedure-execution) @@ -1357,6 +1358,90 @@ nodes: ``` +## Mount Filesystems Procedure + +The `mount_fs` procedure configures filesystems on cluster nodes based on the supplied `procedure.yaml`. + +For each node that contains at least one applicable **fsmount** entry, the behaviour depends on the value of the +``reboot`` option: + +**When ``reboot: true`` (the default):** +1. Drains the node (if it is a ``control-plane`` or ``worker``) to safely evacuate workloads. +2. Removes any existing data from each configured mount path, ensuring a clean state. +3. Installs and enables the corresponding systemd mount units. +4. Reboots the node so the new mounts become active at the operating‑system level. +5. Uncordons the node (again, for ``control-plane`` or ``worker``) to make it schedulable. + +**When ``reboot: false``:** +1. Removes existing data from each configured mount path. +2. Installs and enables the systemd mount units immediately, without performing a reboot. + +Nodes that have no applicable items are skipped entirely. After a successful execution, ``cluster.yaml`` is updated +with the fsmount items defined in the procedure file. + +**Note**: Data inside the configured mount paths is erased before the mount is set up. Back up any important data +prior to running this procedure. +**Note**: For the case when ZRAM volume is mounted to `/var/log/pods`, the kubelet service must be reconfigured +with the recomended parameters. The `procedure.yaml` is as follows: + +```yaml +services: + kubeadm_kubelet: + containerLogMaxSize: 5Mi + containerLogMaxFiles: 2 +``` + +**Warning**: Pay attention, the OS must provide `zram` module. + +### Mount Filesystems Procedure Parameters + +The procedure requires a positional argument that points to the procedure inventory file. + +The JSON schema for the inventory is available at +[URL](/kubemarine/resources/schemas/mount_fs.json?raw=1). See +[Validation by JSON Schemas](Installation.md#inventory-validation) for further details. + +#### ``reboot`` Parameter + +Controls whether each node is drained and rebooted after the mount units are installed. + +* ``true`` (default) – drain, install, reboot, then uncordon. Use this when the filesystem must be cleanly initialised before any + workloads run on the node. +* ``false`` – install and enable the units in‑place without a reboot. Use this when the filesystem can be activated live. + +#### ``fsmount`` Parameter + +The list of filesystem mount items to apply. Its structure is identical to the ``services.fsmount`` section in ``cluster.yaml``. +For a description of the available fields, refer to the [fsmount](Installation.md#fsmount) section in the _Kubemarine Installation Procedure_. + +**Example:** + +```yaml +reboot: true +fsmount: + - name: zram + size: 1G + type: ext4 + device: /dev/zram0 + path: /var/log/pods + template: + source: templates/zram-setup.service.j2 + destination: /etc/systemd/system/zram-setup.service + preparation_scripts: + - resources/scripts/upgrade_kernel.sh + - resources/scripts/zram.sh + groups: [control-plane, worker] +``` + +The node is rebooted between `upgrade_kernel.sh` and `zram.sh` so the new kernel is active before the zram module is loaded. + +### Mount Filesystems Procedure Tasks Tree + +The ``mount_fs`` procedure executes the following sequence of tasks: + +* ``mount_filesystems`` +* ``overview`` + ## Certificate Renew Procedure The `cert_renew` procedure allows you to renew some certificates on an existing Kubernetes cluster. diff --git a/kubemarine/__main__.py b/kubemarine/__main__.py index 29b1781ab..a26942260 100755 --- a/kubemarine/__main__.py +++ b/kubemarine/__main__.py @@ -50,8 +50,8 @@ if (1, 1) in ir[0] and 'tag:yaml.org,2002:float' in ir[1]: float_patched_resolver = (ir[1], ir[2], ir[3]) # Globally change behaviour of yaml.safe_load and yaml.dump - yaml.Dumper.add_implicit_resolver(*float_patched_resolver) # type: ignore[no-untyped-call] - yaml.SafeLoader.add_implicit_resolver(*float_patched_resolver) # type: ignore[no-untyped-call] + yaml.Dumper.add_implicit_resolver(*float_patched_resolver) # type: ignore[no-untyped-call, unused-ignore] + yaml.SafeLoader.add_implicit_resolver(*float_patched_resolver) # type: ignore[no-untyped-call, unused-ignore] break @@ -97,6 +97,10 @@ 'description': "Renew certificates on Kubernetes cluster", 'group': 'maintenance' }, + 'mount_fs': { + 'description': "Mount filesystems described in the fsmount inventory section", + 'group': 'maintenance' + }, 'reboot': { 'description': "Reboot Kubernetes nodes", 'group': 'maintenance' diff --git a/kubemarine/core/resources.py b/kubemarine/core/resources.py index 175fb12fd..f53b31aa9 100644 --- a/kubemarine/core/resources.py +++ b/kubemarine/core/resources.py @@ -33,6 +33,7 @@ import kubemarine.keepalived import kubemarine.kubernetes import kubemarine.kubernetes_accounts +import kubemarine.fsmount import kubemarine.modprobe import kubemarine.packages import kubemarine.plugins @@ -408,6 +409,7 @@ def enrichment_functions(self) -> List[c.EnrichmentFunction]: kubemarine.cri.enrich_upgrade_inventory, kubemarine.plugins.nginx_ingress.cert_renew_enrichment, kubemarine.plugins.envoy_gateway.cert_renew_enrichment, + kubemarine.fsmount.enrich_procedure_inventory, kubemarine.sysctl.enrich_reconfigure_inventory, kubemarine.core.inventory.enrich_reconfigure_inventory, # Enrichment of procedure inventory should be finished at this step. @@ -485,6 +487,7 @@ def enrichment_functions(self) -> List[c.EnrichmentFunction]: kubemarine.system.verify_inventory, kubemarine.system.enrich_etc_hosts, kubemarine.modprobe.enrich_kernel_modules, + kubemarine.fsmount.enrich_inventory, # Calculate some differences between previous and new inventory # Depends on kubemarine.packages.enrich_inventory diff --git a/kubemarine/fsmount.py b/kubemarine/fsmount.py new file mode 100644 index 000000000..e37dfc251 --- /dev/null +++ b/kubemarine/fsmount.py @@ -0,0 +1,226 @@ +# Copyright 2021-2022 NetCracker Technology Corporation +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +import io +import os +from typing import List, Union + +from jinja2 import Template + +from kubemarine.core import utils +from kubemarine.core.cluster import KubernetesCluster, EnrichmentStage, enrichment +from kubemarine.core.group import NodeGroup, CollectorCallback + + +@enrichment(EnrichmentStage.PROCEDURE, procedures=['mount_fs']) +def enrich_procedure_inventory(cluster: KubernetesCluster) -> None: + proc_items: List[dict] = cluster.procedure_inventory.get('fsmount', []) + if not proc_items: + return + + existing: List[dict] = cluster.inventory.setdefault('services', {}).setdefault('fsmount', []) + existing_by_name = {item['name']: i for i, item in enumerate(existing)} + + for item in proc_items: + idx = existing_by_name.get(item['name']) + if idx is not None: + existing[idx] = utils.deepcopy_yaml(item) + else: + existing.append(utils.deepcopy_yaml(item)) + existing_by_name[item['name']] = len(existing) - 1 + + +@enrichment(EnrichmentStage.FULL) +def enrich_inventory(cluster: KubernetesCluster) -> None: + fsmount_list: List[dict] = cluster.inventory.get('services', {}).get('fsmount', []) + for i, item in enumerate(fsmount_list): + path: List[Union[str, int]] = ['services', 'fsmount', i] + + if item.get('groups') is None and item.get('nodes') is None: + continue + + for j, script in enumerate(item.get('preparation_scripts') or []): + ext_path = utils.get_external_resource_path(script) + if not os.path.isfile(ext_path) and not os.path.isfile(utils.get_internal_resource_path(script)): + raise Exception( + f"'preparation_scripts[{j}]' file {script!r} not found " + f"for fsmount item at {utils.pretty_path(path)}") + + if item.get('nodes') is not None: + all_nodes_names = cluster.nodes['all'].get_nodes_names() + unknown_nodes = set(item['nodes']) - set(all_nodes_names) + if unknown_nodes: + cluster.log.warning( + f"Unknown node names {', '.join(map(repr, unknown_nodes))} " + f"provided for fsmount item {item['name']!r}.") + + +def get_applicable_items(cluster: KubernetesCluster, node: NodeGroup, + fsmount_list: List[dict] = None) -> List[dict]: + if fsmount_list is None: + fsmount_list = cluster.inventory.get('services', {}).get('fsmount', []) + applicable = [] + for item in fsmount_list: + groups: Union[List[str], None] = item.get('groups') + nodes: Union[List[str], None] = item.get('nodes') + group = cluster.create_group_from_groups_nodes_names(groups or [], nodes or []) + if group.has_node(node.get_node_name()): + applicable.append(item) + return applicable + + +def _render_unit(item: dict) -> str: + template_source = item['template']['source'] + ext_path = utils.get_external_resource_path(template_source) + if os.path.isfile(ext_path): + template_content = utils.read_external(template_source) + else: + template_content = utils.read_internal(template_source) + + return Template(template_content).render( + name=item['name'], + device=item['device'], + path=item['path'], + size=item.get('size', ''), + type=item.get('type', ''), + ) + + +def _parse_mounts(mounts_output: str) -> dict: + """Parse /proc/mounts into {mountpoint: fstype}.""" + result = {} + for line in mounts_output.splitlines(): + parts = line.split() + if len(parts) >= 3: + result[parts[1]] = parts[2] + return result + + +def is_mounted(group: NodeGroup, fsmount_list: List[dict] = None) -> bool: + cluster: KubernetesCluster = group.cluster + results = group.sudo("cat /proc/mounts") + + for node in group.get_ordered_members_list(): + applicable = get_applicable_items(cluster, node, fsmount_list) + if not applicable: + continue + host = node.get_host() + mounts = _parse_mounts(results[host].stdout) + for item in applicable: + if item['path'].rstrip('/') not in mounts: + cluster.log.debug(f"Mount path {item['path']!r} not found in /proc/mounts on {host}") + return False + + return True + + +def check_mounts(group: NodeGroup, fsmount_list: List[dict] = None) -> List[str]: + """Return a list of human-readable error strings for missing or wrong-type mounts.""" + cluster: KubernetesCluster = group.cluster + + mounts_collector = CollectorCallback(cluster) + zramctl_collector = CollectorCallback(cluster) + defer = group.new_defer() + defer.sudo("cat /proc/mounts", callback=mounts_collector) + defer.sudo("zramctl --output-all", warn=True, callback=zramctl_collector) + defer.flush() + + errors = [] + + for node in group.get_ordered_members_list(): + applicable = get_applicable_items(cluster, node, fsmount_list) + if not applicable: + continue + host = node.get_host() + node_name = node.get_node_name() + mounts = _parse_mounts(mounts_collector.result[host].stdout) + zramctl_output = zramctl_collector.result[host].stdout + + for item in applicable: + mount_path = item['path'].rstrip('/') + expected_type = item.get('type', '') + if mount_path not in mounts: + errors.append(f"{node_name}: {mount_path!r} is not mounted") + elif expected_type and mounts[mount_path] != expected_type: + errors.append( + f"{node_name}: {mount_path!r} has fstype {mounts[mount_path]!r}, expected {expected_type!r}") + + if item['device'].startswith('/dev/zram') and mount_path not in zramctl_output: + errors.append(f"{node_name}: {mount_path!r} not found in zramctl output") + + return errors + + +def setup_fsmount(group: NodeGroup, fsmount_list: List[dict] = None) -> bool: + cluster: KubernetesCluster = group.cluster + logger = cluster.log + + if is_mounted(group, fsmount_list): + logger.debug("Skipped - all required filesystems are already mounted") + return False + + changed = False + for node in group.get_ordered_members_list(): + applicable = get_applicable_items(cluster, node, fsmount_list) + if not applicable: + continue + + host = node.get_host() + mounts_output = node.sudo("cat /proc/mounts")[host].stdout + + for item in applicable: + if item['path'].rstrip('/') in mounts_output: + logger.debug(f"Skipping fsmount item {item['name']!r} on {node.get_node_name()}: already mounted") + continue + + preparation_scripts = item.get('preparation_scripts') or [] + for idx, script_path in enumerate(preparation_scripts): + logger.debug(f"Running preparation script [{idx}] {script_path!r} " + f"for fsmount item {item['name']!r} on {node.get_node_name()}") + ext_path = utils.get_external_resource_path(script_path) + if os.path.isfile(ext_path): + script_content = utils.read_external(script_path) + else: + script_content = utils.read_internal(script_path) + remote_path = f"/tmp/fsmount_{item['name']}_prep_{idx}.sh" + node.put(io.StringIO(script_content), remote_path, sudo=True) + node.sudo(f"chmod +x {remote_path}") + prep_result = node.sudo(f"bash {remote_path}", warn=True) + node.sudo(f"rm -f {remote_path}") + if prep_result[host].return_code != 0: + raise Exception( + f"Preparation script [{idx}] {script_path!r} for fsmount item {item['name']!r} " + f"failed on {node.get_node_name()}. Output: {prep_result[host]}") + + if idx < len(preparation_scripts) - 1: + logger.debug(f"Rebooting {node.get_node_name()!r} between preparation scripts") + initial_boot_history = node.sudo('uptime -s') + node.sudo(cluster.globals['nodes']['boot']['reboot_command'], warn=True) + logger.debug("Waiting for boot up...") + node.wait_for_reboot(initial_boot_history) + + unit_content = _render_unit(item) + unit_destination = item['template']['destination'] + unit_name = unit_destination.rsplit('/', 1)[-1] + unit_dir = unit_destination.rsplit('/', 1)[0] + + logger.debug(f"Setting up fsmount item {item['name']!r} on {node.get_node_name()}") + node.sudo(f"mkdir -p {unit_dir}") + node.put(io.StringIO(unit_content), unit_destination, backup=True, sudo=True) + utils.dump_file(cluster, unit_content, f'fsmount/{item["name"]}_{node.get_node_name()}.service') + node.sudo("systemctl daemon-reload") + node.sudo(f"systemctl enable --now {unit_name}") + changed = True + + return changed diff --git a/kubemarine/procedures/__init__.py b/kubemarine/procedures/__init__.py index 60ee5b593..4fe390735 100644 --- a/kubemarine/procedures/__init__.py +++ b/kubemarine/procedures/__init__.py @@ -53,6 +53,8 @@ def _import_tasks_procedure(name: str) -> TasksProcedure: from kubemarine.procedures import install as procedure elif name == "manage_pss": from kubemarine.procedures import manage_pss as procedure + elif name == "mount_fs": + from kubemarine.procedures import mount_fs as procedure elif name == "reboot": from kubemarine.procedures import reboot as procedure elif name == "reconfigure": diff --git a/kubemarine/procedures/check_paas.py b/kubemarine/procedures/check_paas.py index 9dd6b3148..ee059e083 100755 --- a/kubemarine/procedures/check_paas.py +++ b/kubemarine/procedures/check_paas.py @@ -29,7 +29,7 @@ from kubemarine import ( packages as pckgs, system, selinux, etcd, thirdparties, apparmor, kubernetes, sysctl, audit, - plugins, modprobe, admission + plugins, modprobe, admission, fsmount ) from kubemarine.core.cluster import KubernetesCluster from kubemarine.core.group import NodeGroup, CollectorCallback, GroupResultException @@ -1033,6 +1033,18 @@ def verify_modprobe_rules(cluster: KubernetesCluster) -> None: f"the differences manually and make changes on the appropriate nodes.") +def verify_fsmount(cluster: KubernetesCluster) -> None: + with TestCase(cluster, '236', "System", "Filesystem mounts") as tc: + group = cluster.make_group_from_roles(['control-plane', 'worker']) + errors = fsmount.check_mounts(group) + if not errors: + tc.success(results='mounted') + else: + raise TestFailure('invalid', + hint="Filesystem mount issues found:\n" + "\n".join(f" - {e}" for e in errors) + + "\nRun the fsmount task in the installation procedure to set them up.") + + def verify_sysctl_config(cluster: KubernetesCluster) -> None: """ This test compares the kernel parameters on the nodes @@ -1593,7 +1605,7 @@ def verify_kubernetes_version(cluster: KubernetesCluster) -> None: """ The method checks if used kubernetes version is deprecated in kubemarine """ - with TestCase(cluster, '225', "Kubernetes", "Version") as tc: + with TestCase(cluster, '235', "Kubernetes", "Version") as tc: target_version = cluster.inventory['services']['kubeadm']['kubernetesVersion'] if not kubernetes.verify_supported_version(target_version, cluster.log): raise TestWarn(f"Kubernetes version {target_version} is deprecated", @@ -1763,6 +1775,9 @@ def verify_apparmor_config(cluster: KubernetesCluster) -> None: 'modprobe': { 'rules': verify_modprobe_rules }, + 'fsmount': { + 'mounts': verify_fsmount + }, 'sysctl': { 'config': verify_sysctl_config }, diff --git a/kubemarine/procedures/install.py b/kubemarine/procedures/install.py index 7e0d34cc7..c7a6f265b 100755 --- a/kubemarine/procedures/install.py +++ b/kubemarine/procedures/install.py @@ -22,7 +22,7 @@ from kubemarine.core.errors import KME from kubemarine import ( system, sysctl, haproxy, keepalived, kubernetes, plugins, - kubernetes_accounts, selinux, thirdparties, audit, coredns, cri, packages, apparmor, modprobe + kubernetes_accounts, selinux, thirdparties, audit, coredns, cri, packages, apparmor, modprobe, fsmount ) from kubemarine.core import flow, utils, summary from kubemarine.core.group import NodeGroup, RunnersGroupResult, CollectorCallback @@ -140,6 +140,15 @@ def system_prepare_system_sysctl(group: NodeGroup) -> None: group.call(system.verify_sysctl) +@_applicable_for_new_nodes_with_roles('all') +def system_prepare_system_fsmount(group: NodeGroup) -> None: + cluster: KubernetesCluster = group.cluster + if not cluster.inventory.get('services', {}).get('fsmount'): + cluster.log.debug("Skipped - no fsmount items defined in config file") + return + group.call(fsmount.setup_fsmount) + + @_applicable_for_new_nodes_with_roles('all') def system_prepare_system_setup_selinux(group: NodeGroup) -> None: system.configure_sensitive_service(group, selinux.setup_selinux) @@ -535,6 +544,7 @@ def overview(cluster: KubernetesCluster) -> None: "disable_swap": system_prepare_system_disable_swap, "modprobe": system_prepare_system_modprobe, "sysctl": system_prepare_system_sysctl, + "fsmount": system_prepare_system_fsmount, "audit": { "install": system_install_audit, "configure": system_prepare_audit, @@ -577,9 +587,11 @@ def overview(cluster: KubernetesCluster) -> None: # This is done before `prepare.system.audit`. system.reboot_nodes: [ "prepare.system.modprobe", + "prepare.system.fsmount", "prepare.system.audit" ], system.verify_system: [ + "prepare.system.fsmount", "prepare.system.audit" ], # Some checks can be done only at the end when the necessary services are configured. diff --git a/kubemarine/procedures/mount_fs.py b/kubemarine/procedures/mount_fs.py new file mode 100644 index 000000000..97aa12151 --- /dev/null +++ b/kubemarine/procedures/mount_fs.py @@ -0,0 +1,103 @@ +# Copyright 2021-2022 NetCracker Technology Corporation +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +from collections import OrderedDict +from typing import List + +from kubemarine import fsmount, kubernetes, system +from kubemarine.core import flow +from kubemarine.core.cluster import KubernetesCluster +from kubemarine.procedures import install + + +def mount_filesystems(cluster: KubernetesCluster) -> None: + proc_inv = cluster.procedure_inventory + fsmount_list: List[dict] = proc_inv.get('fsmount', []) + do_reboot: bool = proc_inv.get('reboot', True) + + first_control_plane = cluster.nodes['control-plane'].get_first_member() + timeout_config = cluster.inventory['globals']['expect']['pods']['kubernetes'] + + target_nodes = cluster.make_group([]) + for item in fsmount_list: + target_nodes = target_nodes.include_group( + cluster.create_group_from_groups_nodes_names( + item.get('groups') or [], item.get('nodes') or [])) + + for node in target_nodes.get_ordered_members_list(): + node_name = node.get_node_name() + node_config = node.get_config() + is_k8s_node = 'control-plane' in node_config['roles'] or 'worker' in node_config['roles'] + + applicable = fsmount.get_applicable_items(cluster, node, fsmount_list) + if not applicable: + continue + + if do_reboot and is_k8s_node: + cluster.log.debug(f"Draining node {node_name!r} before fsmount setup") + first_control_plane.sudo( + kubernetes.prepare_drain_command(cluster, node_name, disable_eviction=False), + warn=True, pty=True) + + for item in applicable: + mount_path = item['path'].rstrip('/') + cluster.log.debug(f"Removing files in {mount_path!r} on {node_name!r}") + node.sudo(f"rm -rf {mount_path}/*") + + node.call(fsmount.setup_fsmount, fsmount_list=fsmount_list) + + if do_reboot: + cluster.log.debug(f"Rebooting node {node_name!r} after fsmount setup") + system.perform_group_reboot(node) + + if is_k8s_node: + cluster.log.debug(f"Uncordoning node {node_name!r} after reboot") + first_control_plane.wait_command_successful( + f"kubectl uncordon {node_name}", + hide=False, pty=True, + timeout=timeout_config['timeout'], + retries=timeout_config['retries']) + + +tasks = OrderedDict({ + "mount_filesystems": mount_filesystems, + "overview": install.overview, +}) + + +class MountFsAction(flow.TasksAction): + def __init__(self) -> None: + super().__init__('mount_fs', tasks, recreate_inventory=True) + + +def create_context(cli_arguments: List[str] = None) -> dict: + cli_help = ''' + Script for mounting filesystems defined in procedure.yaml. + + How to use: + + ''' + + parser = flow.new_procedure_parser(cli_help, tasks=tasks) + context = flow.create_context(parser, cli_arguments, procedure='mount_fs') + return context + + +def main(cli_arguments: List[str] = None) -> None: + context = create_context(cli_arguments) + flow.ActionsFlow([MountFsAction()]).run_flow(context) + + +if __name__ == '__main__': + main() diff --git a/kubemarine/resources/configurations/defaults.yaml b/kubemarine/resources/configurations/defaults.yaml index 3367ffc64..a65d84128 100644 --- a/kubemarine/resources/configurations/defaults.yaml +++ b/kubemarine/resources/configurations/defaults.yaml @@ -163,6 +163,8 @@ services: debian: *modprobe-default-modules ubuntu26.04: *modprobe-default-modules + fsmount: [] + sysctl: net.ipv4.ip_nonlocal_bind: value: 1 diff --git a/kubemarine/resources/schemas/definitions/services.json b/kubemarine/resources/schemas/definitions/services.json index cb616a7e6..b44e6410b 100644 --- a/kubemarine/resources/schemas/definitions/services.json +++ b/kubemarine/resources/schemas/definitions/services.json @@ -42,6 +42,9 @@ "modprobe": { "$ref": "services/modprobe.json" }, + "fsmount": { + "$ref": "services/fsmount.json" + }, "sysctl": { "$ref": "services/sysctl.json" }, diff --git a/kubemarine/resources/schemas/definitions/services/fsmount.json b/kubemarine/resources/schemas/definitions/services/fsmount.json new file mode 100644 index 000000000..6e1b297a3 --- /dev/null +++ b/kubemarine/resources/schemas/definitions/services/fsmount.json @@ -0,0 +1,70 @@ +{ + "$schema": "http://json-schema.org/draft-07/schema", + "type": "array", + "description": "List of filesystem mount configurations to be set up on cluster nodes via systemd units.", + "items": { + "oneOf": [ + {"$ref": "#/definitions/FsmountItem"}, + {"$ref": "../common/utils.json#/definitions/ListMergingSymbol"} + ] + }, + "definitions": { + "FsmountItem": { + "type": "object", + "properties": { + "name": { + "type": "string", + "description": "Name of the fsmount item, used for identification" + }, + "device": { + "type": "string", + "description": "The device file to mount (e.g. /dev/sdd1, /dev/zram0)" + }, + "path": { + "type": "string", + "description": "The mount point path on the node" + }, + "size": { + "type": "string", + "description": "Size of the filesystem (optional, e.g. 1G)" + }, + "type": { + "type": "string", + "description": "Filesystem type (e.g. ext4, tmpfs)" + }, + "template": { + "type": "object", + "description": "Systemd unit template configuration", + "properties": { + "source": { + "type": "string", + "description": "Path to the Jinja2 template for the systemd unit" + }, + "destination": { + "type": "string", + "description": "Absolute path on the node where the rendered unit file is placed" + } + }, + "required": ["source", "destination"], + "additionalProperties": false + }, + "preparation_scripts": { + "type": "array", + "description": "Ordered list of shell scripts run before creating the filesystem. A reboot is performed between consecutive scripts. If any script fails, the procedure fails.", + "items": {"type": "string"} + }, + "groups": { + "$ref": "../common/node_ref.json#/definitions/Roles", + "default": ["worker", "control-plane", "balancer"], + "description": "The list of node roles where this mount should be applied" + }, + "nodes": { + "$ref": "../common/node_ref.json#/definitions/Names", + "description": "The list of node names where this mount should be applied" + } + }, + "required": ["name", "device", "path", "template"], + "additionalProperties": false + } + } +} diff --git a/kubemarine/resources/schemas/definitions/services/kubeadm_kubelet.json b/kubemarine/resources/schemas/definitions/services/kubeadm_kubelet.json index 00dc110d9..f9af7e494 100644 --- a/kubemarine/resources/schemas/definitions/services/kubeadm_kubelet.json +++ b/kubemarine/resources/schemas/definitions/services/kubeadm_kubelet.json @@ -10,6 +10,8 @@ "cgroupDriver": {"type": "string", "default": "systemd"}, "maxPods": {"$ref": "#/definitions/MaxPods"}, "serializeImagePulls": {"$ref": "#/definitions/SerializeImagePulls"}, + "containerLogMaxSize": {"$ref": "#/definitions/ContainerLogMaxSize"}, + "containerLogMaxFiles": {"$ref": "#/definitions/ContainerLogMaxFiles"}, "apiVersion": {"type": ["string"], "default": "kubelet.config.k8s.io/v1beta1"}, "kind": {"enum": ["KubeletConfiguration"], "default": "KubeletConfiguration"} }, @@ -17,6 +19,8 @@ "ProtectKernelDefaults": {"type": "boolean", "default": true}, "PodPidsLimit": {"type": "integer", "default": 4096}, "MaxPods": {"type": "integer", "default": 110}, - "SerializeImagePulls": {"type": "boolean", "default": false} + "SerializeImagePulls": {"type": "boolean", "default": false}, + "ContainerLogMaxSize": {"type": "string", "default": "10Mi"}, + "ContainerLogMaxFiles": {"type": "integer", "minimum": 2, "default": 5} } } diff --git a/kubemarine/resources/schemas/mount_fs.json b/kubemarine/resources/schemas/mount_fs.json new file mode 100644 index 000000000..3436bb746 --- /dev/null +++ b/kubemarine/resources/schemas/mount_fs.json @@ -0,0 +1,15 @@ +{ + "$schema": "http://json-schema.org/draft-07/schema", + "type": "object", + "properties": { + "reboot": { + "type": "boolean", + "description": "Whether to drain and reboot the node after mounting. Defaults to true." + }, + "fsmount": { + "$ref": "definitions/services/fsmount.json" + } + }, + "required": ["fsmount"], + "additionalProperties": false +} diff --git a/kubemarine/resources/schemas/reconfigure.json b/kubemarine/resources/schemas/reconfigure.json index 2e0e9ac55..e753fbf14 100644 --- a/kubemarine/resources/schemas/reconfigure.json +++ b/kubemarine/resources/schemas/reconfigure.json @@ -103,7 +103,9 @@ "protectKernelDefaults": {"$ref": "definitions/services/kubeadm_kubelet.json#/definitions/ProtectKernelDefaults"}, "podPidsLimit": {"$ref": "definitions/services/kubeadm_kubelet.json#/definitions/PodPidsLimit"}, "maxPods": {"$ref": "definitions/services/kubeadm_kubelet.json#/definitions/MaxPods"}, - "serializeImagePulls": {"$ref": "definitions/services/kubeadm_kubelet.json#/definitions/SerializeImagePulls"} + "serializeImagePulls": {"$ref": "definitions/services/kubeadm_kubelet.json#/definitions/SerializeImagePulls"}, + "containerLogMaxSize": {"$ref": "definitions/services/kubeadm_kubelet.json#/definitions/ContainerLogMaxSize"}, + "containerLogMaxFiles": {"$ref": "definitions/services/kubeadm_kubelet.json#/definitions/ContainerLogMaxFiles"} }, "additionalProperties": false }, diff --git a/kubemarine/resources/scripts/upgrade_kernel.sh b/kubemarine/resources/scripts/upgrade_kernel.sh new file mode 100644 index 000000000..32518186c --- /dev/null +++ b/kubemarine/resources/scripts/upgrade_kernel.sh @@ -0,0 +1,11 @@ +#!/bin/bash +set -e + +if command -v apt-get &>/dev/null; then + apt-get update -q + apt-get install -yq --only-upgrade linux-image-generic linux-headers-generic +elif command -v dnf &>/dev/null; then + dnf upgrade -y kernel +elif command -v yum &>/dev/null; then + yum upgrade -y kernel +fi diff --git a/kubemarine/resources/scripts/zram.sh b/kubemarine/resources/scripts/zram.sh new file mode 100644 index 000000000..472d8c935 --- /dev/null +++ b/kubemarine/resources/scripts/zram.sh @@ -0,0 +1,11 @@ +#!/bin/bash + +if command -v apt-get &>/dev/null; then + apt-get install -yq linux-modules-extra-$(uname -r) +elif command -v dnf &>/dev/null; then + dnf install -y kernel-modules-extra +elif command -v yum &>/dev/null; then + yum install -y kernel-modules-extra +fi + +mkdir -p /var/log/pods diff --git a/kubemarine/system.py b/kubemarine/system.py index 238f666e3..0fb8b44aa 100644 --- a/kubemarine/system.py +++ b/kubemarine/system.py @@ -23,7 +23,7 @@ from dateutil.parser import parse from ordered_set import OrderedSet -from kubemarine import selinux, apparmor, sysctl, modprobe +from kubemarine import selinux, apparmor, sysctl, modprobe, fsmount from kubemarine.core import utils, static from kubemarine.core.cluster import KubernetesCluster, EnrichmentStage, enrichment from kubemarine.core.executor import RunnersResult, Token, GenericResult, Callback, RawExecutor @@ -571,6 +571,15 @@ def verify_system(cluster: KubernetesCluster) -> None: else: log.debug('Kernel parameters verification skipped - origin setup task was not completed') + if cluster.is_task_completed('prepare.system.fsmount'): + log.debug("Verifying fsmount...") + fsmount_ok = fsmount.is_mounted(group) + if not fsmount_ok: + raise Exception("Required filesystem mounts are not configured") + log.debug("Required filesystem mounts are configured") + else: + log.debug('Fsmount verification skipped - origin setup task was not completed') + def verify_sysctl(group: NodeGroup) -> None: cluster: KubernetesCluster = group.cluster diff --git a/kubemarine/templates/zram-setup.service.j2 b/kubemarine/templates/zram-setup.service.j2 new file mode 100644 index 000000000..5f02066ee --- /dev/null +++ b/kubemarine/templates/zram-setup.service.j2 @@ -0,0 +1,15 @@ +[Unit] +Description=Create zram device +DefaultDependencies=no +Before=local-fs.target + +[Service] +Type=oneshot +RemainAfterExit=yes +ExecStart=/usr/sbin/modprobe zram +ExecStart=/usr/sbin/zramctl {{ device }} --size {{ size }} --algorithm zstd +ExecStart=/usr/sbin/mkfs.{{ type }} -m 0 -O ^has_journal -F {{ device }} +ExecStart=/usr/bin/mount {{ device }} {{ path }} + +[Install] +WantedBy=local-fs.target