This is an automated email from the ASF dual-hosted git repository.

rzo1 pushed a commit to branch security-advisory-stance
in repository https://gitbox.apache.org/repos/asf/stormcrawler-site.git

commit 26dbad6ce2389cb2f465a96a76d4571887b21b5b
Author: Richard Zowalla <[email protected]>
AuthorDate: Tue Aug 25 20:40:32 2026 +0200

    Say when a defect leads to an advisory and when it is hardening
    
    The security model page named a class of defect it treats as a security
    problem, but not what follows from that. This adds the missing half, so a
    report can be triaged consistently:
    
    * the shipped configuration is a starting point, not a secured deployment,
      so a weakness reachable because a documented control was not configured
      is a weak default and the fix is hardening;
    * a control an operator did configure being defeated, or a configured
      secret being disclosed in a way this page does not describe, is a
      vulnerability;
    * an advisory follows when operators must act beyond taking the next
      release, meaning change a configuration, inspect or repair stored data,
      or rotate a credential. Otherwise the fix ships with release notes.
    
    Also states that log output is not a confidentiality boundary, since DEBUG
    logging on the protocol implementations can write configured credentials to
    the logs.
    
    The boundary sentence under "Crawled Content Is Untrusted by Design" now
    says such defects are fixed first and points at this section for whether an
    advisory follows, rather than implying the two are the same question.
---
 security/index.html | 8 +++++++-
 1 file changed, 7 insertions(+), 1 deletion(-)

diff --git a/security/index.html b/security/index.html
index 080dcd8..ea75879 100644
--- a/security/index.html
+++ b/security/index.html
@@ -17,7 +17,7 @@ title: Reporting Security Problems to Apache StormCrawler
 
        <h3>Crawled Content Is Untrusted by Design</h3>
        <p>StormCrawler is a library for broad web crawling. It fetches URLs it 
was told to fetch, including URLs discovered in previously fetched pages, and 
it parses the bytes that come back. Both the URLs and the bytes are chosen by 
the operators of the sites being crawled, which means attacker-controlled input 
is not an edge case for a crawler: it is the normal operating condition. Every 
component that handles a response body, a response header, a redirect target or 
an extracted outlink is  [...]
-       <p>It follows that consequences of the crawler crawling what it was 
configured to crawl are not, in themselves, vulnerabilities. If a topology is 
pointed at a host, fetches a resource from it and stores the result, that is 
the software working as intended, regardless of what the remote host chose to 
serve. What the project does treat as a security problem is content crossing a 
boundary it was never meant to cross: crawled bytes influencing the crawler's 
own configuration or control data [...]
+       <p>It follows that consequences of the crawler crawling what it was 
configured to crawl are not, in themselves, vulnerabilities. If a topology is 
pointed at a host, fetches a resource from it and stores the result, that is 
the software working as intended, regardless of what the remote host chose to 
serve. Content crossing a boundary it was never meant to cross is what the 
project fixes first: crawled bytes influencing the crawler's own configuration 
or control data, reaching resources  [...]
        <p>This is the frame for the rest of this section. The controls 
described below exist so that operators can define where that boundary lies for 
their deployment. Choosing not to configure them widens the boundary rather 
than removing it.</p>
        <p>Malformed or hostile responses that make a crawl slow, stall a 
worker or exhaust its memory are bugs the project fixes. Such effects are 
confined to the crawl the operator chose to run, and the project treats them as 
robustness rather than as a compromise of the deployment.</p>
 
@@ -79,6 +79,12 @@ title: Reporting Security Problems to Apache StormCrawler
                <li>We treat improvements to default configuration values and 
added safeguards as security hardening. These ship in normal releases, 
described in the release notes, without an advisory.</li>
                <li>Vulnerabilities in dependencies follow the ASF's dependency 
guidance. We update affected dependencies in normal releases and do not usually 
issue our own advisory for them.</li>
        </ul>
+       <p>Because StormCrawler is a library rather than a deployed service, 
the configuration it ships is a starting point and not a secured deployment. 
Where a weakness is reachable because a control described on this page was not 
configured, we improve the default and treat the change as hardening. Where the 
library defeats a control an operator did configure, or discloses a configured 
secret in a way this page does not describe, we treat it as a vulnerability.</p>
+
+       <p>We publish an advisory when operators need to act beyond taking the 
next release: change a configuration, inspect or repair stored data, or rotate 
a credential. Where the fix is sufficient on its own, it ships in a normal 
release and the release notes carry any operator-visible change.</p>
+
+       <p>Log output is not a confidentiality boundary. Enabling DEBUG logging 
on the protocol implementations can put configured credentials into the logs. 
Treat DEBUG output from a topology that uses an authenticated proxy or crawls 
authenticated sites as sensitive, and scope log retention accordingly.</p>
+
        <p>As the ASF policy notes, what counts as a vulnerability depends in 
part on what a project states as expected behaviour, which is why this page 
describes the crawler's threat model rather than only its reporting process.</p>
 
        <h3>Summary</h3>

Reply via email to