* test: migrate test_namespace.sh to CI/tests_v2 namespace_deletion Migrate the legacy CI/tests/test_namespace.sh functional test to the pytest-based v2 framework under CI/tests_v2/scenarios/namespace_deletion/. Covers the service_disruption (namespace deletion) scenario with happy-path cases (single-namespace object deletion, multi-namespace delete_count, multiple runs, wait_time handling, label-selector targeting) and negative cases (no-match regex, namespace/label mutual exclusion, delete_count exceeding available namespaces). - Add scenario assets (resource.yaml with Deployment+Service, scenario_base.yaml) - Register namespace_deletion marker in pytest.ini - Add namespace_deletion execution-evidence marker in lib/utils.py Closes #1405 * test: move namespace_deletion helpers into reusable lib modules Address review feedback: extract the per-test helpers into shared lib modules so other scenarios can reuse them. - lib/namespace.py: add POD_SECURITY_PRIVILEGED_LABELS, create_labeled_namespace, delete_namespace_quietly, and a make_namespace factory fixture (auto-cleanup) - lib/deploy.py: add deployment_exists, wait_for_no_deployment, wait_for_present_deployment_count, deploy_manifest_to_namespace - conftest.py: re-export make_namespace fixture - test_namespace_deletion.py: drop local helpers, use the lib functions/fixture * test: honor --keep-ns-on-fail in make_namespace factory The make_namespace finalizer always deleted ad-hoc namespaces, ignoring --keep-ns-on-fail on failure unlike test_namespace. Extract the keep-on-fail decision into a shared helper and use it in both fixtures so the documented debugging workflow works for scenarios that create extra namespaces. * test: surface FailToCreateError in deploy_manifest_to_namespace Mirror deploy_workload's error handling so manifest-apply failures in multi-namespace tests raise a formatted RuntimeError listing the underlying API exceptions instead of an opaque FailToCreateError. * test: clarify why label-selector test bypasses run_scenario The inline comment claimed run_scenario was bypassed because **overrides would collide with the positional namespace arg. The real reason is that NAMESPACE_IS_REGEX=True wraps an empty namespace as '^$', whereas label-selector mode needs a literal empty string. * test: use actual scenario inputs in negative-test failure context The no-match and mutual-exclusion tests reported context=namespace=self.ns (the ephemeral namespace), not the inputs actually under test. Reference the real namespace (and label_selector) so unexpected-success diagnostics are clear. * test: use non-default wait_time so override patching is exercised wait_time=30 matched the scenario_base.yaml default, so the test passed even if override patching regressed. Use wait_time=5 (non-default) so the override path is actually validated. * test: guard cluster post-checks under KRKN_TEST_DRY_RUN Three namespace_deletion tests asserted cluster side effects (workload deletion) after the Kraken run. Under KRKN_TEST_DRY_RUN=1 Kraken is skipped, so the seeded workload is never deleted and wait_for_no_deployment / wait_for_present_deployment_count would time out and fail. Guard those post-checks, and in test_label_selector_targeting (which bypasses run_scenario and calls run_kraken directly) honor dry-run explicitly by skipping the invocation and post-check. * test: tighten no-match assertion, clarify runs-loop test, dry-run-safe negatives Addresses Deep Code Review feedback on the namespace_deletion suite: - test_no_match_namespace_fails: drop the dead 'no namespaces matching' OR branch; the service_disruption plugin only ever logs 'not enough namespaces matching ...', so the assertion now checks that string directly. - test_multiple_runs_repeat_deletion -> test_multiple_runs_repeat_disruption_loop: rename + docstring make explicit that it verifies the outer runs loop iterates twice, not that object deletion recurs (Krkn does not redeploy between runs, so run 2 re-selects an already-empty namespace). Object removal is asserted in test_single_namespace_object_deletion. - Negative tests now return early under KRKN_TEST_DRY_RUN=1, since run_scenario returns a fake rc=0 and the failure path cannot be exercised; this makes the whole class consistent under make test-dry-run. * test: use UUID-based namespace in no-match test to avoid accidental matches * test: assert correct zero-match error in no-match namespace test A regex matching zero namespaces makes krkn_lib's check_namespaces raise 'there exists no namespaces matching' before the plugin's delete loop, so the 'not enough namespaces matching' branch is never reached for this case. Assert the actual zero-match message instead. * test: poll for async Service deletion in namespace_deletion test Kubernetes deletions are asynchronous, so checking the Service immediately after the scenario run could be flaky. Add wait_for_no_service (mirroring wait_for_no_deployment) and poll until the Service is actually gone before asserting. * test: assert Service deletion in label-selector namespace_deletion test test_label_selector_targeting deploys both a Deployment and a Service but only asserted the Deployment was removed. Add wait_for_no_service so the test fully validates that label-selector mode deletes all objects (including Services), matching test_single_namespace_object_deletion. Signed-off-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com> * test: snapshot namespaces in wait_for_present_deployment_count list(namespaces) consumed the iterable before the poll loop re-iterated over the original, so a one-shot iterable (e.g. generator) would be empty on every poll. Snapshot to a list once and iterate over that snapshot. Signed-off-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com> * ci: free up runner disk space before tests_v2 KinD run The Tests v2 (pytest functional) job intermittently fails with 'System.IO.IOException: No space left on device' on ubuntu-latest runners, which ship with only ~14GB free. Creating the KinD cluster plus pulling and kind-loading nginx:alpine and krkn:tools exhausts the disk, failing even the runner's own diagnostic logging. Reclaim ~20-30GB by removing the bundled .NET/Android/GHC SDKs and pruning docker images before the cluster is created. Signed-off-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com> --------- Signed-off-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com> Co-authored-by: augmentcode[bot] <185243770+augmentcode[bot]@users.noreply.github.com> Co-authored-by: Darshan Jain <darjain@redhat.com>
Krkn aka Kraken
Chaos and resiliency testing tool for Kubernetes. Kraken injects deliberate failures into Kubernetes clusters to check if it is resilient to turbulent conditions.
Workflow
How to Get Started
Instructions on how to setup, configure and run Kraken can be found in the documentation.
Blogs, podcasts and interviews
Additional resources, including blog posts, podcasts, and community interviews, can be found on the website
Roadmap
Enhancements being planned can be found in the roadmap.
Contributing
We are always looking for more enhancements, fixes to make it better, any contributions are most welcome. Feel free to report or work on the issues filed on github.
See CONTRIBUTING.md for details, or check the contribution guidelines on the website.
Community
Key Members(slack_usernames/full name): paigerube14/Paige Rubendall, mffiedler/Mike Fiedler, tsebasti/Tullio Sebastiani, yogi/Yogananth Subramanian, sahil/Sahil Shah, pradeep/Pradeep Surisetty and ravielluri/Naga Ravi Chaitanya Elluri.
Community Meeting
If you have any questions that you think could be better discussed on a meeting, we have monthly office hours. Please add items to the agenda beforehand so we can best prepare to help you.
The Linux Foundation® (TLF) has registered trademarks and uses trademarks. For a list of TLF trademarks, see Trademark Usage.

