kratos0718 opened a new pull request, #73617: URL: https://github.com/apache/airflow/pull/73617
_mask_cmd masked passwords/secrets in the spark-submit command line using a regex with a lazy prefix that gets re-attempted from every position in the string. A user-controlled `application_arg` containing neither "secret" nor "password" made this quadratic in that argument's length — a single ~50k character argument tied up the call for close to a minute, and an authenticated user with DAG-trigger permissions could repeat this to exhaust worker slots. Replaced the regex with token-by-token scanning: split the command on whitespace, check each token with a plain substring test, and only do the (still linear) quote-handling work for tokens that actually contain "secret" or "password". Verified this reproduces the exact same masked output as the old regex for every existing test case, including a multi-word quoted value spread across several whitespace-separated tokens — a case the straightforward "just check each token" approach breaks unless you re-merge tokens on an unclosed quote, which is what `_mask_value` does. closes: #70676 ## Test Plan Ran `breeze testing providers-tests providers/apache/spark/tests/unit/apache/spark/hooks/test_spark_submit.py -k mask` — 12 passed (the 8 existing cases unchanged, 2 new cases for the multi-word-quote edge case, 1 new timing regression test). Also verified independently, outside the test suite, that the fix stays linear (all under ~2ms) at up to 1.28M characters across several adversarial shapes: no secret/password anywhere in the string, secret/password present but buried in a large noisy token, multiple occurrences mixed with noise, and an attacker-opened quote that's never closed across many tokens. The original regex was quadratic on the first of these (confirmed independently: 0.19s / 0.72s / 2.92s at 5k/10k/20k chars) and several of the others. `ruff format` / `ruff check --fix`, `prek run --from-ref main --stage pre-commit`, `prek run --from-ref main --stage manual` (with the doc-server/skill-eval hooks skipped), and `breeze run mypy` on the changed file all pass clean. --- ##### Was generative AI tooling used to co-author this PR? - [X] Yes (please specify the tool below) Generated-by: Claude Code (Sonnet 5) following [the guidelines](https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst#gen-ai-assisted-contributions) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
