Files
krkn/CI/tests_v2/README.md
T
augmentcode[bot]andlnx01 03416ddbb6 test: migrate node_scenarios to tests_v2 pytest framework (#1400)
* test: add node_scenarios pytest migration under tests_v2

Migrate the node chaos coverage from the legacy CI/tests/test_node.sh into the v2 pytest framework at CI/tests_v2/scenarios/node_scenarios/.

- test_node_scenarios.py (TestNodeScenarios extends BaseScenarioTest): reboot and stop/start happy paths, node_name vs label_selector targeting, node recovery with finalizer, a control-plane safety guard, and negative cases (invalid selector, invalid node, missing actions, unsupported cloud type, unknown action).

- scenario_base.yaml: single node_scenarios entry (cloud_type docker, worker-only) patched per test.

- Register the node_scenarios marker in pytest.ini and document the scenario in CI/tests_v2/README.md.

Part of #1398. The coupled legacy move + workflow edit is left for a maintainer (the bot lacks GitHub App workflows permission); see PR description.

* test: address node_scenarios review feedback

- Reboot happy path now runs with kube_check: True so Krkn waits for the
  node to go Unknown then Ready, eliminating the race where wait_node_ready
  could pass against a stale Ready=True before the disruption propagated.
- Finalizer ensure_node_container_running now polls for Ready (bounded) when
  given k8s_core and logs non-zero 'start' exits, matching the docstring
  contract so a rerun never picks up an unrecovered node.
- Clarify the parallelism note: first/last worker separation only applies on
  multi-worker clusters (CI's 2-worker kind-config.yml); single-worker dev
  clusters share the node and rely on test ordering.

* fix: make node happy-path tests resilient to KinD multi-CP API LB

The reboot/stop_start happy paths failed in the Tests v2 job with
"Response ended prematurely": with kube_check enabled, Krkn's docker node
plugin polls the kube API (wait_for_unknown_status/wait_for_ready_status)
for ~40-50s after the disruption, and krkn-lib does not retry a transient
connection drop from the multi-control-plane KinD haproxy API load balancer.
The docker action itself succeeded in <0.5s.

Run both happy paths with kube_check: False so Krkn performs the action and
exits cleanly, and prove real disruption deterministically via the node
container's State.StartedAt advancing (runtime-level evidence, independent of
node-status timing). Recovery is still verified with the resilient test-side
wait_node_ready poll. README updated to match.

* refactor: move reusable node/container test helpers to lib/utils

Per review feedback, relocate the generic node-level helpers out of the
node_scenarios test module into the shared CI/tests_v2/lib/utils.py so future
node/container tests can reuse them: wait_node_ready, container_runtime,
container_started_at, assert_container_cycled, ensure_node_container_running,
and assert_kraken_marker. The test module now imports them; behavior is
unchanged and all 8 tests still collect.

* docs: fix duplicated word in tests_v2 README scenario list

* test: skip redundant container start in node finalizer to avoid noisy warnings

* test: honor KIND_EXPERIMENTAL_PROVIDER when selecting container runtime

* docs: align node finalizer wording with start-if-stopped behavior

* docs: align node test module docstring with start-if-stopped finalizer

* test: sort schedulable worker nodes for deterministic targeting

* test: assert combined node_stop_start_scenario marker in stop/start happy path

---------

Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
2026-06-15 14:51:15 +05:30

17 KiB

Pytest Functional Tests (tests_v2)

This directory contains a pytest-based functional test framework that runs alongside the existing bash tests in CI/tests/. It covers the pod disruption, application outage, storage throttle, CPU hog, memory hog, and node scenarios with proper assertions, retries, and reporting.

Each test runs in its own ephemeral Kubernetes namespace (krkn-test-<uuid>). Before the test, the framework creates the namespace, deploys the target workload, and waits for pods to be ready. After the test, the namespace is deleted (cascading all resources). You do not need to deploy any workloads manually.

Prerequisites

Without a cluster, tests that need one will skip with a clear message (e.g. "Could not load kube config"). No manual workload deployment is required; workloads are deployed automatically into ephemeral namespaces per test.

  • KinD cluster (or any Kubernetes cluster) running with kubectl configured (e.g. KUBECONFIG or default ~/.kube/config).
  • Python 3.11+ and main repo deps: pip install -r requirements.txt.

Supported clusters

  • KinD (recommended): Use make -f CI/tests_v2/Makefile setup from the repo root. Fastest for local dev; uses a 2-node dev config by default. Override with KIND_CONFIG=/path/to/kind-config.yml for a larger cluster.
  • Minikube: Should work; ensure kubectl context is set. Not tested in CI.
  • Remote/cloud cluster: Tests create and delete namespaces; use with caution. Use --require-kind to avoid accidentally running against production (tests will skip unless context is kind/minikube).

Setting up the cluster

Option A: Use the setup script (recommended)

From the repository root, with kind and kubectl installed:

# Create KinD cluster (defaults to CI/tests_v2/kind-config-dev.yml; override with KIND_CONFIG=...)
./CI/tests_v2/setup_env.sh

Then in the same shell (or after export KUBECONFIG=~/.kube/config in another terminal), activate your venv and install Python deps:

python3 -m venv venv
source venv/bin/activate   # or: source venv/Scripts/activate on Windows
pip install -r requirements.txt
pip install -r CI/tests_v2/requirements.txt

Option B: Manual setup

  1. Install kind and kubectl.
  2. Create a cluster (from repo root):
    kind create cluster --name kind --config kind-config.yml
    
  3. Wait for the cluster:
    kubectl wait --for=condition=Ready nodes --all --timeout=120s
    
  4. Create a virtualenv, activate it, and install dependencies (as in Option A).
  5. Run tests from repo root: pytest CI/tests_v2/ -v ...

Install test dependencies

From the repository root:

pip install -r CI/tests_v2/requirements.txt

This adds pytest-rerunfailures, pytest-html, pytest-timeout, and pytest-order (pytest and coverage come from the main requirements.txt).

Dependency Management

Dependencies are split into two files:

  • Root requirements.txt — Kraken runtime (cloud SDKs, Kubernetes client, krkn-lib, pytest, coverage, etc.). Required to run Kraken.
  • CI/tests_v2/requirements.txt — Test-only pytest plugins (rerunfailures, html, timeout, order, xdist). Not needed by Kraken itself.

Rule of thumb: If Kraken needs it at runtime, add to root. If only the functional tests need it, add to CI/tests_v2/requirements.txt.

Running make -f CI/tests_v2/Makefile setup (or make setup from CI/tests_v2) creates the venv and installs both files automatically; you do not need to install them separately. The Makefile re-installs when either file changes (via the .installed sentinel).

Run tests

All commands below are from the repository root.

Basic run (with retries and HTML report)

pytest CI/tests_v2/ -v --timeout=300 --reruns=2 --reruns-delay=10 --html=CI/tests_v2/report.html --junitxml=CI/tests_v2/results.xml
  • Failed tests are retried up to 2 times with a 10s delay (configurable in CI/tests_v2/pytest.ini).
  • Each test has a 5-minute timeout.
  • Open CI/tests_v2/report.html in a browser for a detailed report.

Run in parallel (faster suite)

pytest CI/tests_v2/ -v -n 4 --timeout=300

Ephemeral namespaces make tests parallel-safe; use -n with the number of workers (e.g. 4).

Run without retries (for debugging)

pytest CI/tests_v2/ -v -p no:rerunfailures

Run with coverage

python -m coverage run -m pytest CI/tests_v2/ -v
python -m coverage report

To append to existing coverage from unit tests, ensure coverage was started with coverage run -a for earlier runs, or run the full test suite in one go.

Run only pod disruption tests

pytest CI/tests_v2/ -v -m pod_disruption

Run only application outage tests

pytest CI/tests_v2/ -v -m application_outage

Run only CPU hog tests

pytest CI/tests_v2/ -v -m cpu_hog

Run only memory hog tests

pytest CI/tests_v2/ -v -m memory_hog

Run only node scenarios tests

pytest CI/tests_v2/ -v -m node_scenarios

Note: Node scenarios stop/reboot a worker node (a KinD node is a Docker/Podman container), which is cluster-wide disruption. Run them on a KinD cluster you can afford to disrupt, and prefer a multi-worker cluster when combining with other scenarios under -n auto (see the parallelism note in scenarios/node_scenarios/test_node_scenarios.py).

Run with verbose output and no capture

pytest CI/tests_v2/ -v -s

Keep failed test namespaces for debugging

When a test fails, its ephemeral namespace is normally deleted. To keep the namespace so you can inspect pods, logs, and network policies:

pytest CI/tests_v2/ -v --keep-ns-on-fail

On failure, the namespace name is printed (e.g. [keep-ns-on-fail] Keeping namespace krkn-test-a1b2c3d4 for debugging). Use kubectl get pods -n krkn-test-a1b2c3d4 (and similar) to debug, then delete the namespace manually when done.

Logging and cluster options

  • Structured logging: Use --log-cli-level=DEBUG to see namespace creation, workload deploy, and readiness in the console. Use --log-file=test.log to capture logs to a file.
  • Require dev cluster: To avoid running against the wrong cluster, use --require-kind. Tests will skip unless the current kube context cluster name contains "kind" or "minikube".
  • Stale namespace cleanup: At session start, namespaces matching krkn-test-* that are older than 30 minutes are deleted (e.g. from a previous crashed run).
  • Timeout overrides: Set env vars to tune timeouts (e.g. in CI): KRKN_TEST_READINESS_TIMEOUT, KRKN_TEST_DEPLOY_TIMEOUT, KRKN_TEST_NS_CLEANUP_TIMEOUT, KRKN_TEST_POLICY_WAIT_TIMEOUT, KRKN_TEST_KRAKEN_PROC_WAIT_TIMEOUT, KRKN_TEST_TIMEOUT_BUDGET.

Architecture

  • Folder-per-scenario: Each scenario lives under scenarios/<scenario_name>/ with:
    • test_.py — Test class extending BaseScenarioTest; sets WORKLOAD_MANIFEST, SCENARIO_NAME, SCENARIO_TYPE, NAMESPACE_KEY_PATH, and optionally OVERRIDES_KEY_PATH.
    • resource.yaml — Kubernetes resources (Deployment/Pod) for the scenario; namespace is patched at deploy time.
    • scenario_base.yaml — Canonical Krkn scenario; the base class loads it, patches namespace (and overrides), and passes it to Kraken via run_scenario(). Optional extra YAMLs (e.g. nginx_http.yaml for application_outage) can live in the same folder.
  • lib/: Shared framework — lib/base.py defines BaseScenarioTest, timeout constants (env-overridable), and scenario helpers (load_and_patch_scenario, run_scenario); lib/utils.py provides assertion and K8s helpers; lib/k8s.py provides K8s client fixtures; lib/namespace.py provides namespace lifecycle; lib/deploy.py provides deploy_workload, wait_for_pods_running, wait_for_deployment_replicas; lib/kraken.py provides run_kraken, build_config (using CI/tests_v2/config/common_test_config.yaml).
  • conftest.py: Re-exports fixtures from the lib modules and defines pytest_addoption, logging, and repo_root.
  • Adding a new scenario: Use the scaffold script (see CONTRIBUTING_TESTS.md) to create scenarios/<name>/ with test file, resource.yaml, and scenario_base.yaml, or copy an existing scenario folder and adapt.

What is tested

Each test runs in an isolated ephemeral namespace; workloads are deployed automatically before the test and the namespace is deleted after (unless --keep-ns-on-fail is set and the test failed).

  • scenarios/pod_disruption/
    Pod disruption scenario. resource.yaml is a deployment with label app=krkn-pod-disruption-target; scenario_base.yaml is loaded and namespace_pattern is patched to the test namespace. The test:

    1. Records baseline pod UIDs and restart counts.
    2. Runs Kraken with the pod disruption scenario.
    3. Asserts that chaos had an effect (UIDs changed or restart count increased).
    4. Waits for pods to be Running and all containers Ready.
    5. Asserts pod count is unchanged and all pods are healthy.
  • scenarios/application_outage/
    Application outage scenario (block Ingress/Egress to target pods, then restore). resource.yaml is the main workload (outage pod); scenario_base.yaml is loaded and patched with namespace (and duration/block as needed). Optional nginx_http.yaml is used by the traffic test. Tests include:

    • test_app_outage_block_restore_and_variants: Happy path with default, exclude_label, and block variants (Ingress, Egress, both); Krkn exit 0, pods still Running/Ready.
    • test_network_policy_created_then_deleted: Policy with prefix krkn-deny- appears during run and is gone after.
    • test_traffic_blocked_during_outage (disabled, planned): Deploys nginx with label scenario=outage, port-forwards; during outage curl fails, after run curl succeeds.
    • test_invalid_scenario_fails: Invalid scenario file (missing application_outage key) causes Kraken to exit non-zero.
    • test_bad_namespace_fails: Scenario targeting a non-existent namespace causes Kraken to exit non-zero.
  • scenarios/cpu_hog/ CPU hog scenario (hog_scenarios), migrated from the legacy CI/tests/test_cpu_hog.sh. CPU hog targets nodes (not workloads): Kraken deploys a short-lived hog pod (name prefix cpu-hog-) onto each selected node, runs stress-ng for the configured duration, then deletes the pod. Tests use @pytest.mark.no_workload (no app deployment needed); scenario_base.yaml is a flat hog config patched per test. Tests include:

    • test_cpu_hog_success_lifecycle_and_targeting: Happy path — a hog pod is created on the node-selector target, the run exits 0, and the pod is cleaned up afterward.
    • test_cpu_hog_invalid_selector_fails: A node-selector matching zero nodes causes Kraken to exit non-zero (no available nodes to schedule).
    • test_cpu_hog_invalid_config_fails: Omitting the mandatory hog-type field causes Kraken to exit non-zero at config parsing.
  • scenarios/memory_hog/ Memory hog scenario (hog_scenarios), migrated from the legacy CI/tests/test_memory_hog.sh. Memory hog targets nodes (not workloads): Kraken deploys a short-lived hog pod (name prefix memory-hog-) onto each selected node, runs stress-ng for the configured duration with the configured memory-vm-bytes, then deletes the pod. Tests use @pytest.mark.no_workload (no app deployment needed); scenario_base.yaml is a flat hog config patched per test (with a small fixed memory-vm-bytes instead of the production 90%). Tests include:

    • test_memory_hog_success_lifecycle_and_targeting: Happy path — a hog pod is created on the node-selector target with the configured memory size, the run exits 0, and the pod is cleaned up afterward.
    • test_memory_hog_invalid_selector_fails: A node-selector matching zero nodes causes Kraken to exit non-zero (no available nodes to schedule).
    • test_memory_hog_invalid_config_fails: Omitting the mandatory hog-type field causes Kraken to exit non-zero at config parsing.
  • scenarios/node_scenarios/ Node chaos scenario (node_scenarios), migrated from the legacy CI/tests/test_node.sh. Node scenarios are destructive at the node level: Kraken stops, starts, or reboots the container that backs a Kubernetes node. On KinD each node is a Docker/Podman container, so the tests use cloud_type: docker and target a worker node only (never the control plane). Tests use @pytest.mark.no_workload (no app deployment needed); scenario_base.yaml holds a single node_scenarios entry that each test patches per case. A finalizer ensures the targeted node container is running (starting it only if it was left stopped) and waits for the node to return Ready. Tests include:

    • test_node_reboot_targets_node_name_and_recovers: Happy path — node_reboot_scenario targeted by node_name reboots the worker. Kraken runs with kube_check: False (its in-process Unknown→Ready wait is brittle behind the multi-control-plane KinD API load balancer); disruption is proven by the node container's StartedAt advancing and recovery by the node returning Ready.
    • test_node_stop_start_targets_label_selector_and_recovers: Happy path — node_stop_start_scenario targeted by label_selector stops then starts the worker. Same approach: kube_check: False, disruption proven by StartedAt advancing and recovery by the node returning Ready.
    • test_invalid_label_selector_fails: A label_selector matching zero nodes causes Kraken to exit non-zero.
    • test_invalid_node_name_fails: A non-existent node_name (not a killable node) causes Kraken to exit non-zero.
    • test_missing_actions_fails: An entry without actions causes Kraken to fail fast.
    • test_unsupported_cloud_type_fails: An unsupported cloud_type causes Kraken to exit non-zero when building the node scenario object.
    • test_unknown_action_is_skipped: An unrecognized action is skipped (logged, no node touched) and Kraken still exits 0.
    • test_control_plane_excluded_from_targeting: Safety guard — control-plane/master nodes are never returned by the worker-targeting helper the destructive tests use.

Configuration

  • pytest.ini: Markers (functional, pod_disruption, application_outage, storage_throttle, cpu_hog, memory_hog, node_scenarios, no_workload). Use --timeout=300, --reruns=2, --reruns-delay=10 on the command line for full runs.
  • conftest.py: Re-exports fixtures from lib/k8s.py, lib/namespace.py, lib/deploy.py, lib/kraken.py (e.g. test_namespace, deploy_workload, k8s_core, wait_for_pods_running, run_kraken, build_config). Configs are built from CI/tests_v2/config/common_test_config.yaml with monitoring disabled for local runs. Timeout constants in lib/base.py can be overridden via env vars.
  • Cluster access: Reads and applies use the Kubernetes Python client; kubectl is still used for port-forward and for running Kraken.
  • utils.py: Pod/network policy helpers and assertion helpers (assert_all_pods_running_and_ready, assert_pod_count_unchanged, assert_kraken_success, assert_kraken_failure, patch_namespace_in_docs).

Relationship to existing CI

  • The existing bash tests in CI/tests/ and CI/run.sh are unchanged. They continue to run as before in GitHub Actions.
  • This framework is additive. To run it in CI later, add a separate job or step that runs pytest CI/tests_v2/ ... from the repo root.

Troubleshooting

  • pytest.skip: Could not load kube config — No cluster or bad KUBECONFIG. Run make -f CI/tests_v2/Makefile setup (or make setup from CI/tests_v2) or check kubectl cluster-info.
  • KinD cluster creation hangs — Docker is not running. Start Docker Desktop or run systemctl start docker.
  • Bind for 0.0.0.0:9090 failed: port is already allocated — Another process (e.g. Prometheus) is using the port. The default dev config (kind-config-dev.yml) no longer maps host ports; if you use KIND_CONFIG=kind-config.yml or a custom config with extraPortMappings, free the port or switch to kind-config-dev.yml.
  • TimeoutError: Pods did not become ready — Slow image pull or node resource limits. Increase KRKN_TEST_READINESS_TIMEOUT or check node resources.
  • ModuleNotFoundError: pytest_rerunfailures — Missing test deps. Run pip install -r CI/tests_v2/requirements.txt (or make setup).
  • Stale krkn-test-* namespaces — Left over from a previous crashed run. They are auto-cleaned at session start (older than 30 min). To remove cluster and reports: make -f CI/tests_v2/Makefile clean.
  • Wrong cluster targeted — Multiple kube contexts. Use --require-kind to skip unless context is kind/minikube, or set context explicitly: kubectl config use-context kind-ci-krkn.
  • OSError: [Errno 48] Address already in use when running tests in parallel — Kraken normally starts an HTTP status server on port 8081. With -n auto (pytest-xdist), multiple Kraken processes would all try to bind to 8081. The test framework disables this server (publish_kraken_status: False) in the generated config, so parallel runs should not hit this. If you see it, ensure you're using the framework's build_config and not a config that has publish_kraken_status: True.