This is an automated email from the ASF dual-hosted git repository. rzo1 pushed a commit to branch security-advisory-stance in repository https://gitbox.apache.org/repos/asf/stormcrawler-site.git
commit 26dbad6ce2389cb2f465a96a76d4571887b21b5b Author: Richard Zowalla <[email protected]> AuthorDate: Tue Aug 25 20:40:32 2026 +0200 Say when a defect leads to an advisory and when it is hardening The security model page named a class of defect it treats as a security problem, but not what follows from that. This adds the missing half, so a report can be triaged consistently: * the shipped configuration is a starting point, not a secured deployment, so a weakness reachable because a documented control was not configured is a weak default and the fix is hardening; * a control an operator did configure being defeated, or a configured secret being disclosed in a way this page does not describe, is a vulnerability; * an advisory follows when operators must act beyond taking the next release, meaning change a configuration, inspect or repair stored data, or rotate a credential. Otherwise the fix ships with release notes. Also states that log output is not a confidentiality boundary, since DEBUG logging on the protocol implementations can write configured credentials to the logs. The boundary sentence under "Crawled Content Is Untrusted by Design" now says such defects are fixed first and points at this section for whether an advisory follows, rather than implying the two are the same question. --- security/index.html | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/security/index.html b/security/index.html index 080dcd8..ea75879 100644 --- a/security/index.html +++ b/security/index.html @@ -17,7 +17,7 @@ title: Reporting Security Problems to Apache StormCrawler <h3>Crawled Content Is Untrusted by Design</h3> <p>StormCrawler is a library for broad web crawling. It fetches URLs it was told to fetch, including URLs discovered in previously fetched pages, and it parses the bytes that come back. Both the URLs and the bytes are chosen by the operators of the sites being crawled, which means attacker-controlled input is not an edge case for a crawler: it is the normal operating condition. Every component that handles a response body, a response header, a redirect target or an extracted outlink is [...] - <p>It follows that consequences of the crawler crawling what it was configured to crawl are not, in themselves, vulnerabilities. If a topology is pointed at a host, fetches a resource from it and stores the result, that is the software working as intended, regardless of what the remote host chose to serve. What the project does treat as a security problem is content crossing a boundary it was never meant to cross: crawled bytes influencing the crawler's own configuration or control data [...] + <p>It follows that consequences of the crawler crawling what it was configured to crawl are not, in themselves, vulnerabilities. If a topology is pointed at a host, fetches a resource from it and stores the result, that is the software working as intended, regardless of what the remote host chose to serve. Content crossing a boundary it was never meant to cross is what the project fixes first: crawled bytes influencing the crawler's own configuration or control data, reaching resources [...] <p>This is the frame for the rest of this section. The controls described below exist so that operators can define where that boundary lies for their deployment. Choosing not to configure them widens the boundary rather than removing it.</p> <p>Malformed or hostile responses that make a crawl slow, stall a worker or exhaust its memory are bugs the project fixes. Such effects are confined to the crawl the operator chose to run, and the project treats them as robustness rather than as a compromise of the deployment.</p> @@ -79,6 +79,12 @@ title: Reporting Security Problems to Apache StormCrawler <li>We treat improvements to default configuration values and added safeguards as security hardening. These ship in normal releases, described in the release notes, without an advisory.</li> <li>Vulnerabilities in dependencies follow the ASF's dependency guidance. We update affected dependencies in normal releases and do not usually issue our own advisory for them.</li> </ul> + <p>Because StormCrawler is a library rather than a deployed service, the configuration it ships is a starting point and not a secured deployment. Where a weakness is reachable because a control described on this page was not configured, we improve the default and treat the change as hardening. Where the library defeats a control an operator did configure, or discloses a configured secret in a way this page does not describe, we treat it as a vulnerability.</p> + + <p>We publish an advisory when operators need to act beyond taking the next release: change a configuration, inspect or repair stored data, or rotate a credential. Where the fix is sufficient on its own, it ships in a normal release and the release notes carry any operator-visible change.</p> + + <p>Log output is not a confidentiality boundary. Enabling DEBUG logging on the protocol implementations can put configured credentials into the logs. Treat DEBUG output from a topology that uses an authenticated proxy or crawls authenticated sites as sensitive, and scope log retention accordingly.</p> + <p>As the ASF policy notes, what counts as a vulnerability depends in part on what a project states as expected behaviour, which is why this page describes the crawler's threat model rather than only its reporting process.</p> <h3>Summary</h3>
