AlKor13 commented on issue #73525: URL: https://github.com/apache/airflow/issues/73525#issuecomment-5847425867
Still reproducible on `main` @ `f863982`, and the collection path has the same defect in two more places than the one you named — so fixing the script alone would move the failure rather than remove it. `scripts/ci/prek/update_providers_dependencies.py`, five `Path.read_text()` / `write_text()` calls with no `encoding=`: ``` L73 tomllib.loads(pyproject_toml_file_path.read_text()) L98 yaml.safe_load(provider_yaml_file.read_text()) L230 DEPENDENCIES_JSON_FILE_PATH.read_text() ... L233 old_content = DEPENDENCIES_JSON_FILE_PATH.read_text() ... L235 DEPENDENCIES_JSON_FILE_PATH.write_text(new_content) ``` `devel-common/src/tests_common/pytest_plugin.py`, which is what invokes it during collection, three more: ``` L169 AIRFLOW_PYPROJECT_TOML_FILE_PATH.read_text().splitlines() L201 PROVIDER_DEPENDENCIES_JSON_HASH_PATH.read_text() L203 PROVIDER_DEPENDENCIES_JSON_HASH_PATH.write_text(calculated_provider_deps_hash) ``` L169 reads the root `pyproject.toml` in the locale code page before the script is ever called, so on cp1252 a single non-ASCII byte anywhere in that file fails collection with the same `UnicodeDecodeError` even after the script is fixed. L235/L203 are the write side: on a non-UTF-8 machine they write `generated/provider_dependencies.json` and its hash file in the ANSI code page, which is then committed or compared against a UTF-8 original. A static pass over `scripts/`, `dev/`, `devel-common/` and `airflow-core/src/` (sparse checkout, so not the whole repo) finds **1,102** places of this shape: 844 `read_text`/`write_text`, 166 `open()` in text mode, 77 `subprocess(text=True)`, 15 `shutil.rmtree()` without `onexc=`, concentrated in `scripts/` (557) and `dev/` (455). Two notes on scoping the fix: **A repo-wide `encoding="utf-8"` is not busywork that 3.15 will undo.** [PEP 686](https://peps.python.org/pep-0686/) makes UTF-8 the default in Python 3.15, but Airflow supports 3.10+, so every user below 3.15 keeps the bug; naming the encoding is what the 3.15 migration would leave behind anyway. **The 77 `subprocess(text=True)` call sites are a different fix from the file ones.** Windows has two default code pages at once — on this machine cp1251 (what `text=True` decodes with) and cp866 (what a console child writes) — so UTF-8 is correct for a Python child and wrong for `git` or `uv`; after 3.15 those turn from silent mojibake into an exception. Disclosure: the inventory came from `winseam audit`, a tool I wrote after hitting this repeatedly ([repo](https://github.com/AlKor13/winseam)); it is a static pass, so it runs on Linux CI. Happy to send a PR for the eight lines in the collection path if that is useful — that is the part that unblocks Windows test runs. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
