exec_cmd_in_pod wraps commands with `bash -c` when no base_command is
specified. Passing ["ip", "-br", "addr", "show"] produces
["bash", "-c", "ip", "-br", "addr", "show"]. Due to bash -c semantics,
only the first argument after -c ("ip") is treated as the command
string; the rest become unused positional parameters. This caused bare
`ip` to run with no arguments, printing help/usage text instead of
interface data.
The egress scenario was unaffected because it uses base_command="chroot"
which constructs commands correctly without bash -c wrapping.
Pass each command as a single string (e.g. ["ip -br addr show"]) so
bash -c treats it as one complete command.
Closes#1380
Signed-off-by: Rahul Shetty <rashetty@redhat.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* feat: migrate container scenarios tests to pytest v2
Migrate CI/tests/test_container.sh to CI/tests_v2/scenarios/container_scenarios/.
Adds dry_run support to ContainerScenarioPlugin and moves legacy test to CI/legacy/.
Closes#1399
Signed-off-by: qsxDree <kurling.town@gmail.com>
* test: strengthen container label selector and dry-run behavior
Deploy a decoy workload to verify label_selector filtering, and skip pod
monitoring when dry_run is enabled so dry runs return without waiting for
expected_recovery_time.
Signed-off-by: qsxDree <kurling.town@gmail.com>
* drop dry_run and keep v1 container test
Signed-off-by: qsxDree <kurling.town@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix container scenario test assertions
Signed-off-by: qsxDree <kurling.town@gmail.com>
---------
Signed-off-by: qsxDree <kurling.town@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Darshan Jain <darjain@redhat.com>
krkn-lib 6.1.0 declared kubernetes>=34.1.0 (unbounded), allowing pip to
resolve to kubernetes 36.x which renamed call_api() parameter
response_type to response_types_map, breaking SA token creation, service
patching, and node metrics queries.
krkn-lib 6.1.2 pins kubernetes>=35.0.0,<36.0.0, preventing the
incompatible version from being installed.
Fixes#1441
Signed-off-by: ddjain <darjain@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: validate container_name in container scenario plugin
The container scenario plugin removed pods and incremented killed_count
even when the requested container_name was never found, causing scenarios
to silently report success with no disruption (issue #1409).
Track whether a container was actually found and killed; only increment
killed_count on a real kill, skip pods that lack the target container,
and raise a clear RuntimeError once all pods are exhausted without a kill.
Adds unit tests covering invalid, valid, empty, heterogeneous, and
count-exceeds-target scenarios.
Closes#1409
* fix: report actual kill count in container-not-found error
When the candidate pod list is exhausted without finding the target
container, the RuntimeError now reports how many containers were
actually killed ("N of M requested container(s) were killed") instead
of always claiming "No containers were killed", which was inaccurate in
partial-success cases.
* fix: only raise container-not-found error when nothing was killed
Address review feedback: the "not found in any matching pod" error was
raised even after one or more containers had already been killed (when
count exceeds the number of pods containing the target), making the
message contradictory.
Now that error only fires when killed_count == 0. When some kills
already happened but the candidate list is exhausted, the loop falls
through to the existing "Trying to kill more containers than were found"
error, which accurately describes that case.
---------
Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
Co-authored-by: Darshan Jain <darjain@redhat.com>
* test: migrate test_namespace.sh to CI/tests_v2 namespace_deletion
Migrate the legacy CI/tests/test_namespace.sh functional test to the pytest-based
v2 framework under CI/tests_v2/scenarios/namespace_deletion/.
Covers the service_disruption (namespace deletion) scenario with happy-path cases
(single-namespace object deletion, multi-namespace delete_count, multiple runs,
wait_time handling, label-selector targeting) and negative cases (no-match regex,
namespace/label mutual exclusion, delete_count exceeding available namespaces).
- Add scenario assets (resource.yaml with Deployment+Service, scenario_base.yaml)
- Register namespace_deletion marker in pytest.ini
- Add namespace_deletion execution-evidence marker in lib/utils.py
Closes#1405
* test: move namespace_deletion helpers into reusable lib modules
Address review feedback: extract the per-test helpers into shared lib
modules so other scenarios can reuse them.
- lib/namespace.py: add POD_SECURITY_PRIVILEGED_LABELS, create_labeled_namespace,
delete_namespace_quietly, and a make_namespace factory fixture (auto-cleanup)
- lib/deploy.py: add deployment_exists, wait_for_no_deployment,
wait_for_present_deployment_count, deploy_manifest_to_namespace
- conftest.py: re-export make_namespace fixture
- test_namespace_deletion.py: drop local helpers, use the lib functions/fixture
* test: honor --keep-ns-on-fail in make_namespace factory
The make_namespace finalizer always deleted ad-hoc namespaces, ignoring
--keep-ns-on-fail on failure unlike test_namespace. Extract the keep-on-fail
decision into a shared helper and use it in both fixtures so the documented
debugging workflow works for scenarios that create extra namespaces.
* test: surface FailToCreateError in deploy_manifest_to_namespace
Mirror deploy_workload's error handling so manifest-apply failures in
multi-namespace tests raise a formatted RuntimeError listing the underlying
API exceptions instead of an opaque FailToCreateError.
* test: clarify why label-selector test bypasses run_scenario
The inline comment claimed run_scenario was bypassed because **overrides
would collide with the positional namespace arg. The real reason is that
NAMESPACE_IS_REGEX=True wraps an empty namespace as '^$', whereas
label-selector mode needs a literal empty string.
* test: use actual scenario inputs in negative-test failure context
The no-match and mutual-exclusion tests reported context=namespace=self.ns
(the ephemeral namespace), not the inputs actually under test. Reference the
real namespace (and label_selector) so unexpected-success diagnostics are clear.
* test: use non-default wait_time so override patching is exercised
wait_time=30 matched the scenario_base.yaml default, so the test passed even
if override patching regressed. Use wait_time=5 (non-default) so the override
path is actually validated.
* test: guard cluster post-checks under KRKN_TEST_DRY_RUN
Three namespace_deletion tests asserted cluster side effects (workload
deletion) after the Kraken run. Under KRKN_TEST_DRY_RUN=1 Kraken is
skipped, so the seeded workload is never deleted and wait_for_no_deployment
/ wait_for_present_deployment_count would time out and fail. Guard those
post-checks, and in test_label_selector_targeting (which bypasses
run_scenario and calls run_kraken directly) honor dry-run explicitly by
skipping the invocation and post-check.
* test: tighten no-match assertion, clarify runs-loop test, dry-run-safe negatives
Addresses Deep Code Review feedback on the namespace_deletion suite:
- test_no_match_namespace_fails: drop the dead 'no namespaces matching' OR
branch; the service_disruption plugin only ever logs 'not enough namespaces
matching ...', so the assertion now checks that string directly.
- test_multiple_runs_repeat_deletion -> test_multiple_runs_repeat_disruption_loop:
rename + docstring make explicit that it verifies the outer runs loop iterates
twice, not that object deletion recurs (Krkn does not redeploy between runs, so
run 2 re-selects an already-empty namespace). Object removal is asserted in
test_single_namespace_object_deletion.
- Negative tests now return early under KRKN_TEST_DRY_RUN=1, since run_scenario
returns a fake rc=0 and the failure path cannot be exercised; this makes the
whole class consistent under make test-dry-run.
* test: use UUID-based namespace in no-match test to avoid accidental matches
* test: assert correct zero-match error in no-match namespace test
A regex matching zero namespaces makes krkn_lib's check_namespaces raise
'there exists no namespaces matching' before the plugin's delete loop, so
the 'not enough namespaces matching' branch is never reached for this case.
Assert the actual zero-match message instead.
* test: poll for async Service deletion in namespace_deletion test
Kubernetes deletions are asynchronous, so checking the Service immediately
after the scenario run could be flaky. Add wait_for_no_service (mirroring
wait_for_no_deployment) and poll until the Service is actually gone before
asserting.
* test: assert Service deletion in label-selector namespace_deletion test
test_label_selector_targeting deploys both a Deployment and a Service but
only asserted the Deployment was removed. Add wait_for_no_service so the
test fully validates that label-selector mode deletes all objects
(including Services), matching test_single_namespace_object_deletion.
Signed-off-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
* test: snapshot namespaces in wait_for_present_deployment_count
list(namespaces) consumed the iterable before the poll loop re-iterated
over the original, so a one-shot iterable (e.g. generator) would be empty
on every poll. Snapshot to a list once and iterate over that snapshot.
Signed-off-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
* ci: free up runner disk space before tests_v2 KinD run
The Tests v2 (pytest functional) job intermittently fails with
'System.IO.IOException: No space left on device' on ubuntu-latest runners,
which ship with only ~14GB free. Creating the KinD cluster plus pulling and
kind-loading nginx:alpine and krkn:tools exhausts the disk, failing even the
runner's own diagnostic logging.
Reclaim ~20-30GB by removing the bundled .NET/Android/GHC SDKs and pruning
docker images before the cluster is created.
Signed-off-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
---------
Signed-off-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
Co-authored-by: Darshan Jain <darjain@redhat.com>
The IO hog scenario reported negative disk space values because:
1. disk_space from get_node_resources_info() returns availableBytes
(free space), not used space
2. The formula subtracted from node_resources_end (misleadingly named,
actually measured 3s after pod deploy) instead of node_resources_start
When the IO hog writes data, available space decreases. The original
formula (avg - end) produced negative values. The correct formula is
(start - avg) which measures how much available space was consumed.
Tested on AKS (2-node cluster) with io-write-bytes=512m:
- Before: -512.24 MB, -495.16 MB, -23.72 MB (5/5 negative)
- After: +512.38 MB, +512.23 MB, +500.48 MB (3/3 correct)
Signed-off-by: Shiva <shiva@users.noreply.github.com>
Signed-off-by: SK8-infi <shivansh.katiyar1712@gmail.com>
* test: add node_scenarios pytest migration under tests_v2
Migrate the node chaos coverage from the legacy CI/tests/test_node.sh into the v2 pytest framework at CI/tests_v2/scenarios/node_scenarios/.
- test_node_scenarios.py (TestNodeScenarios extends BaseScenarioTest): reboot and stop/start happy paths, node_name vs label_selector targeting, node recovery with finalizer, a control-plane safety guard, and negative cases (invalid selector, invalid node, missing actions, unsupported cloud type, unknown action).
- scenario_base.yaml: single node_scenarios entry (cloud_type docker, worker-only) patched per test.
- Register the node_scenarios marker in pytest.ini and document the scenario in CI/tests_v2/README.md.
Part of #1398. The coupled legacy move + workflow edit is left for a maintainer (the bot lacks GitHub App workflows permission); see PR description.
* test: address node_scenarios review feedback
- Reboot happy path now runs with kube_check: True so Krkn waits for the
node to go Unknown then Ready, eliminating the race where wait_node_ready
could pass against a stale Ready=True before the disruption propagated.
- Finalizer ensure_node_container_running now polls for Ready (bounded) when
given k8s_core and logs non-zero 'start' exits, matching the docstring
contract so a rerun never picks up an unrecovered node.
- Clarify the parallelism note: first/last worker separation only applies on
multi-worker clusters (CI's 2-worker kind-config.yml); single-worker dev
clusters share the node and rely on test ordering.
* fix: make node happy-path tests resilient to KinD multi-CP API LB
The reboot/stop_start happy paths failed in the Tests v2 job with
"Response ended prematurely": with kube_check enabled, Krkn's docker node
plugin polls the kube API (wait_for_unknown_status/wait_for_ready_status)
for ~40-50s after the disruption, and krkn-lib does not retry a transient
connection drop from the multi-control-plane KinD haproxy API load balancer.
The docker action itself succeeded in <0.5s.
Run both happy paths with kube_check: False so Krkn performs the action and
exits cleanly, and prove real disruption deterministically via the node
container's State.StartedAt advancing (runtime-level evidence, independent of
node-status timing). Recovery is still verified with the resilient test-side
wait_node_ready poll. README updated to match.
* refactor: move reusable node/container test helpers to lib/utils
Per review feedback, relocate the generic node-level helpers out of the
node_scenarios test module into the shared CI/tests_v2/lib/utils.py so future
node/container tests can reuse them: wait_node_ready, container_runtime,
container_started_at, assert_container_cycled, ensure_node_container_running,
and assert_kraken_marker. The test module now imports them; behavior is
unchanged and all 8 tests still collect.
* docs: fix duplicated word in tests_v2 README scenario list
* test: skip redundant container start in node finalizer to avoid noisy warnings
* test: honor KIND_EXPERIMENTAL_PROVIDER when selecting container runtime
* docs: align node finalizer wording with start-if-stopped behavior
* docs: align node test module docstring with start-if-stopped finalizer
* test: sort schedulable worker nodes for deterministic targeting
* test: assert combined node_stop_start_scenario marker in stop/start happy path
---------
Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
* test: migrate memory_hog functional test to tests_v2 (closes#1396)
Add a pytest v2 functional test for the memory hog scenario (hog_scenarios),
mirroring the cpu_hog migration. Covers execution success, node-selector
targeting, duration/memory-size parameter handling, hog pod lifecycle and
cleanup, and graceful failure on invalid selector / invalid config.
- CI/tests_v2/scenarios/memory_hog/test_memory_hog.py: TestMemoryHog with
functional + memory_hog markers and three no_workload tests.
- CI/tests_v2/scenarios/memory_hog/scenario_base.yaml: flat hog config tuned
for functional testing (light fixed memory-vm-bytes, short duration).
- Register the memory_hog marker in pytest.ini and document the scenario in
the tests_v2 README.
* test: extract shared hog-pod/node helpers into tests_v2 lib/utils
Lift the duplicated pod-prefix and schedulable-node helpers out of the cpu_hog and memory_hog test modules into CI/tests_v2/lib/utils.py as parameterized, reusable functions (list_pods_by_prefix, wait_for_scheduled_pod_by_prefix, wait_for_no_pods_by_prefix, schedulable_worker_nodes) and reuse them from both scenarios. Addresses review feedback on #1397.
---------
Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
* test: migrate cpu_hog functional test to tests_v2 (closes#1390)
Migrate the legacy CI/tests/test_cpu_hog.sh to the v2 pytest framework
under CI/tests_v2/scenarios/cpu_hog/, preserving parity with the legacy
flow and adding stronger functional and negative coverage.
- Add scenario_base.yaml (flat hog_scenarios config, patched per test)
- Add test_cpu_hog.py with TestCpuHog(BaseScenarioTest):
- success: hog pod created on node-selector target, run exits 0, pod cleaned up
- negative: node-selector matching zero nodes fails gracefully
- negative: missing mandatory hog-type fails at config parsing
- Register cpu_hog marker in pytest.ini
- Document cpu_hog coverage in README.md
* test: kill background Kraken proc on any cpu_hog success-test failure
Broaden the success-path teardown so a poll/assert failure before
proc.communicate() also kills the background Kraken process, preventing
a lingering cpu-hog- pod from stressing the node and racing --reruns.
Addresses Deep Code Review feedback on PR #1391.
---------
Signed-off-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
* test: attach krkn execution logs to HTML report and assert scenario executed
CI/tests_v2 enhancements for observability and false-positive guarding:
- run_kraken fixture stashes stdout/stderr/returncode on request.node so the
report hook can attach them.
- conftest pytest_runtest_makereport renders the krkn log as a timestamp/level/
message table via pytest_html.extras for every test (pass or fail).
- conftest pytest_terminal_summary prints an execution-evidence table and writes
a markdown summary to $GITHUB_STEP_SUMMARY when running in GitHub Actions.
- utils adds SCENARIO_EXECUTION_MARKERS, an EXECUTION_EVIDENCE registry, and
assert_scenario_executed() that fails happy-path tests when krkn exits 0 but
no scenario-specific marker is present. Skipped under KRKN_TEST_DRY_RUN=1.
- pod_disruption, storage_throttle, and application_outage happy-path tests now
call assert_scenario_executed. Negative tests remain exempt.
Closes#1394
* test: stash krkn logs on timeout and add log-path hint to evidence assert
Address Deep Code Review feedback on PR #1395:
- run_kraken now catches subprocess.TimeoutExpired, stashes partial
stdout/stderr with a synthetic rc=124 so timed-out runs still attach
logs to the HTML report.
- assert_scenario_executed failure message now includes the
'Full logs:' tmp_path hint, matching assert_kraken_success/failure.
---------
Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
- Add exec_with_shell_fallback method with retry logic and shell fallback
- Add unit tests for the new method with proper mocking
- All tests now pass as expected
Signed-off-by: NITESH SINGH <niteshkumar121411@gmail.com>
Co-authored-by: Paige Patton <64206430+paigerube14@users.noreply.github.com>
Scorecard Token-Permissions check (alert #4) flags release.yml for
missing top-level permissions, which means GITHUB_TOKEN defaults to
broad write access across all jobs. Adding permissions: read-all at
the top level enforces least privilege by default; the release job
already declares contents: write at job level for the permissions it
actually needs.
Signed-off-by: Paige Patton <prubenda@redhat.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Allows selecting VMIs by label selector as an alternative to vm_name regex,
making vm_name optional when label_selector is provided.
Signed-off-by: Paige Patton <prubenda@redhat.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
The bare except clause in the Kubernetes client initialization block
caused a NameError crash when initialization failed. If KrknKubernetes()
raised an exception before kubecli was assigned, the except handler
attempted to call kubecli.initialize_clients(None) on an undefined
variable, masking the original error entirely.
Replaced the bare except with except Exception as e to:
- Log the actual initialization error for visibility
- Initialize both kubecli and ocpcli with None kubeconfig as fallback
so subsequent code referencing these variables does not crash
- Avoid catching SystemExit and KeyboardInterrupt unintentionally
Fixes#1264
Signed-off-by: Parth Agrawal <parth.agrawal4002@gmail.com>
Co-authored-by: Paige Patton <64206430+paigerube14@users.noreply.github.com>
The start_klusterlet_scenario branch in inject_managedcluster_scenario
was calling stop_klusterlet_scenario on the scenarios object instead of
start_klusterlet_scenario. Any user configuring this action would stop
the klusterlet (scale to 0) instead of starting it (scale to 3).
Fixes#1323
Signed-off-by: v0idheaven <dahiyavarun2007@gmail.com>
Co-authored-by: Paige Patton <64206430+paigerube14@users.noreply.github.com>
* fix: use .get() for optional health check config keys to prevent KeyError
bearer_token, auth, exit_on_failure and verify_url were accessed with
bare [] indexing. If a user omits any of these optional keys, the health
check thread crashes with a KeyError that disappears silently — the
telemetry queue never gets populated and health checks stop working
with no log output.
Replaced all four with .get() calls with safe defaults.
Fixes#1309
Signed-off-by: v0idheaven <dahiyavarun2007@gmail.com>
* fix: address review feedback on HealthChecker config key handling
- Add docstring to run_health_check for API clarity
- Replace conditional url assignment with direct .get() + continue
so a missing url skips the entry cleanly instead of falling through
to make_request with a potentially stale url from a previous iteration
Signed-off-by: v0idheaven <dahiyavarun2007@gmail.com>
---------
Signed-off-by: v0idheaven <dahiyavarun2007@gmail.com>
Co-authored-by: Paige Patton <64206430+paigerube14@users.noreply.github.com>