* test: add node_scenarios pytest migration under tests_v2
Migrate the node chaos coverage from the legacy CI/tests/test_node.sh into the v2 pytest framework at CI/tests_v2/scenarios/node_scenarios/.
- test_node_scenarios.py (TestNodeScenarios extends BaseScenarioTest): reboot and stop/start happy paths, node_name vs label_selector targeting, node recovery with finalizer, a control-plane safety guard, and negative cases (invalid selector, invalid node, missing actions, unsupported cloud type, unknown action).
- scenario_base.yaml: single node_scenarios entry (cloud_type docker, worker-only) patched per test.
- Register the node_scenarios marker in pytest.ini and document the scenario in CI/tests_v2/README.md.
Part of #1398. The coupled legacy move + workflow edit is left for a maintainer (the bot lacks GitHub App workflows permission); see PR description.
* test: address node_scenarios review feedback
- Reboot happy path now runs with kube_check: True so Krkn waits for the
node to go Unknown then Ready, eliminating the race where wait_node_ready
could pass against a stale Ready=True before the disruption propagated.
- Finalizer ensure_node_container_running now polls for Ready (bounded) when
given k8s_core and logs non-zero 'start' exits, matching the docstring
contract so a rerun never picks up an unrecovered node.
- Clarify the parallelism note: first/last worker separation only applies on
multi-worker clusters (CI's 2-worker kind-config.yml); single-worker dev
clusters share the node and rely on test ordering.
* fix: make node happy-path tests resilient to KinD multi-CP API LB
The reboot/stop_start happy paths failed in the Tests v2 job with
"Response ended prematurely": with kube_check enabled, Krkn's docker node
plugin polls the kube API (wait_for_unknown_status/wait_for_ready_status)
for ~40-50s after the disruption, and krkn-lib does not retry a transient
connection drop from the multi-control-plane KinD haproxy API load balancer.
The docker action itself succeeded in <0.5s.
Run both happy paths with kube_check: False so Krkn performs the action and
exits cleanly, and prove real disruption deterministically via the node
container's State.StartedAt advancing (runtime-level evidence, independent of
node-status timing). Recovery is still verified with the resilient test-side
wait_node_ready poll. README updated to match.
* refactor: move reusable node/container test helpers to lib/utils
Per review feedback, relocate the generic node-level helpers out of the
node_scenarios test module into the shared CI/tests_v2/lib/utils.py so future
node/container tests can reuse them: wait_node_ready, container_runtime,
container_started_at, assert_container_cycled, ensure_node_container_running,
and assert_kraken_marker. The test module now imports them; behavior is
unchanged and all 8 tests still collect.
* docs: fix duplicated word in tests_v2 README scenario list
* test: skip redundant container start in node finalizer to avoid noisy warnings
* test: honor KIND_EXPERIMENTAL_PROVIDER when selecting container runtime
* docs: align node finalizer wording with start-if-stopped behavior
* docs: align node test module docstring with start-if-stopped finalizer
* test: sort schedulable worker nodes for deterministic targeting
* test: assert combined node_stop_start_scenario marker in stop/start happy path
---------
Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
* test: migrate memory_hog functional test to tests_v2 (closes#1396)
Add a pytest v2 functional test for the memory hog scenario (hog_scenarios),
mirroring the cpu_hog migration. Covers execution success, node-selector
targeting, duration/memory-size parameter handling, hog pod lifecycle and
cleanup, and graceful failure on invalid selector / invalid config.
- CI/tests_v2/scenarios/memory_hog/test_memory_hog.py: TestMemoryHog with
functional + memory_hog markers and three no_workload tests.
- CI/tests_v2/scenarios/memory_hog/scenario_base.yaml: flat hog config tuned
for functional testing (light fixed memory-vm-bytes, short duration).
- Register the memory_hog marker in pytest.ini and document the scenario in
the tests_v2 README.
* test: extract shared hog-pod/node helpers into tests_v2 lib/utils
Lift the duplicated pod-prefix and schedulable-node helpers out of the cpu_hog and memory_hog test modules into CI/tests_v2/lib/utils.py as parameterized, reusable functions (list_pods_by_prefix, wait_for_scheduled_pod_by_prefix, wait_for_no_pods_by_prefix, schedulable_worker_nodes) and reuse them from both scenarios. Addresses review feedback on #1397.
---------
Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
* test: migrate cpu_hog functional test to tests_v2 (closes#1390)
Migrate the legacy CI/tests/test_cpu_hog.sh to the v2 pytest framework
under CI/tests_v2/scenarios/cpu_hog/, preserving parity with the legacy
flow and adding stronger functional and negative coverage.
- Add scenario_base.yaml (flat hog_scenarios config, patched per test)
- Add test_cpu_hog.py with TestCpuHog(BaseScenarioTest):
- success: hog pod created on node-selector target, run exits 0, pod cleaned up
- negative: node-selector matching zero nodes fails gracefully
- negative: missing mandatory hog-type fails at config parsing
- Register cpu_hog marker in pytest.ini
- Document cpu_hog coverage in README.md
* test: kill background Kraken proc on any cpu_hog success-test failure
Broaden the success-path teardown so a poll/assert failure before
proc.communicate() also kills the background Kraken process, preventing
a lingering cpu-hog- pod from stressing the node and racing --reruns.
Addresses Deep Code Review feedback on PR #1391.
---------
Signed-off-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
* test: attach krkn execution logs to HTML report and assert scenario executed
CI/tests_v2 enhancements for observability and false-positive guarding:
- run_kraken fixture stashes stdout/stderr/returncode on request.node so the
report hook can attach them.
- conftest pytest_runtest_makereport renders the krkn log as a timestamp/level/
message table via pytest_html.extras for every test (pass or fail).
- conftest pytest_terminal_summary prints an execution-evidence table and writes
a markdown summary to $GITHUB_STEP_SUMMARY when running in GitHub Actions.
- utils adds SCENARIO_EXECUTION_MARKERS, an EXECUTION_EVIDENCE registry, and
assert_scenario_executed() that fails happy-path tests when krkn exits 0 but
no scenario-specific marker is present. Skipped under KRKN_TEST_DRY_RUN=1.
- pod_disruption, storage_throttle, and application_outage happy-path tests now
call assert_scenario_executed. Negative tests remain exempt.
Closes#1394
* test: stash krkn logs on timeout and add log-path hint to evidence assert
Address Deep Code Review feedback on PR #1395:
- run_kraken now catches subprocess.TimeoutExpired, stashes partial
stdout/stderr with a synthetic rc=124 so timed-out runs still attach
logs to the HTML report.
- assert_scenario_executed failure message now includes the
'Full logs:' tmp_path hint, matching assert_kraken_success/failure.
---------
Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com>
- Add exec_with_shell_fallback method with retry logic and shell fallback
- Add unit tests for the new method with proper mocking
- All tests now pass as expected
Signed-off-by: NITESH SINGH <niteshkumar121411@gmail.com>
Co-authored-by: Paige Patton <64206430+paigerube14@users.noreply.github.com>
Scorecard Token-Permissions check (alert #4) flags release.yml for
missing top-level permissions, which means GITHUB_TOKEN defaults to
broad write access across all jobs. Adding permissions: read-all at
the top level enforces least privilege by default; the release job
already declares contents: write at job level for the permissions it
actually needs.
Signed-off-by: Paige Patton <prubenda@redhat.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Allows selecting VMIs by label selector as an alternative to vm_name regex,
making vm_name optional when label_selector is provided.
Signed-off-by: Paige Patton <prubenda@redhat.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
The bare except clause in the Kubernetes client initialization block
caused a NameError crash when initialization failed. If KrknKubernetes()
raised an exception before kubecli was assigned, the except handler
attempted to call kubecli.initialize_clients(None) on an undefined
variable, masking the original error entirely.
Replaced the bare except with except Exception as e to:
- Log the actual initialization error for visibility
- Initialize both kubecli and ocpcli with None kubeconfig as fallback
so subsequent code referencing these variables does not crash
- Avoid catching SystemExit and KeyboardInterrupt unintentionally
Fixes#1264
Signed-off-by: Parth Agrawal <parth.agrawal4002@gmail.com>
Co-authored-by: Paige Patton <64206430+paigerube14@users.noreply.github.com>
The start_klusterlet_scenario branch in inject_managedcluster_scenario
was calling stop_klusterlet_scenario on the scenarios object instead of
start_klusterlet_scenario. Any user configuring this action would stop
the klusterlet (scale to 0) instead of starting it (scale to 3).
Fixes#1323
Signed-off-by: v0idheaven <dahiyavarun2007@gmail.com>
Co-authored-by: Paige Patton <64206430+paigerube14@users.noreply.github.com>
* fix: use .get() for optional health check config keys to prevent KeyError
bearer_token, auth, exit_on_failure and verify_url were accessed with
bare [] indexing. If a user omits any of these optional keys, the health
check thread crashes with a KeyError that disappears silently — the
telemetry queue never gets populated and health checks stop working
with no log output.
Replaced all four with .get() calls with safe defaults.
Fixes#1309
Signed-off-by: v0idheaven <dahiyavarun2007@gmail.com>
* fix: address review feedback on HealthChecker config key handling
- Add docstring to run_health_check for API clarity
- Replace conditional url assignment with direct .get() + continue
so a missing url skips the entry cleanly instead of falling through
to make_request with a potentially stale url from a previous iteration
Signed-off-by: v0idheaven <dahiyavarun2007@gmail.com>
---------
Signed-off-by: v0idheaven <dahiyavarun2007@gmail.com>
Co-authored-by: Paige Patton <64206430+paigerube14@users.noreply.github.com>
* fix(cerberus): fix config key typo and shadowed global in get_status
Fixes#1302
- Fix typo in set_url(): read 'check_application_routes' instead of
'check_applicaton_routes' so the correct config key is used
- Remove local 'check_application_routes = False' in get_status() and
declare global instead, so the value set by set_url() is actually read
- Update test config key and remove the workaround comment that
acknowledged the shadowed global
Signed-off-by: netram75 <netram.24bcs10329@sst.scaler.com>
* fix(cerberus): remove extra cerberus_url arg from application_status call
application_status(start_time, end_time) takes two parameters but was
being called with three (cerberus_url, start_time, end_time), which
would raise a TypeError whenever check_application_routes is enabled.
Remove the stale argument since the function already reads cerberus_url
from the module global.
Signed-off-by: netram75 <netram.24bcs10329@sst.scaler.com>
---------
Signed-off-by: netram75 <netram.24bcs10329@sst.scaler.com>
Co-authored-by: Paige Patton <64206430+paigerube14@users.noreply.github.com>