This is an automated email from the ASF dual-hosted git repository.
rzo1 pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/stormcrawler-site.git
The following commit(s) were added to refs/heads/main by this push:
new e404bed Say when a defect leads to an advisory and when it is
hardening (#46)
e404bed is described below
commit e404bed3faba8a62975765b3b594630412fca7f1
Author: Richard Zowalla <[email protected]>
AuthorDate: Thu Aug 27 13:50:17 2026 +0200
Say when a defect leads to an advisory and when it is hardening (#46)
* Say when a defect leads to an advisory and when it is hardening
The security model page named a class of defect it treats as a security
problem, but not what follows from that. This adds the missing half, so a
report can be triaged consistently:
* the shipped configuration is a starting point, not a secured deployment,
so a weakness reachable because a documented control was not configured
is a weak default and the fix is hardening;
* a control an operator did configure being defeated, or a configured
secret being disclosed in a way this page does not describe, is a
vulnerability;
* an advisory follows when operators must act beyond taking the next
release, meaning change a configuration, inspect or repair stored data,
or rotate a credential. Otherwise the fix ships with release notes.
Also states that log output is not a confidentiality boundary, since DEBUG
logging on the protocol implementations can write configured credentials to
the logs.
The boundary sentence under "Crawled Content Is Untrusted by Design" now
says such defects are fixed first and points at this section for whether an
advisory follows, rather than implying the two are the same question.
* Cover stored data and name availability in the robustness statement
The robustness statement only mentioned hostile responses, so a worker
crash driven by data the crawl stored earlier and replayed did not clearly
fall under it. It now covers both paths, uses the word availability once so
the position is findable, and distinguishes the crawl's own availability
from that of the Storm cluster, which is the operator's to protect.
* Do not make the advisory depend on how much work the fix needs
The bullet list says a vulnerability in a released artifact gets an
advisory. The paragraph added in this branch said we publish one when
operators have to act beyond taking the next release, which contradicts it
and would have left the ordinary case, where upgrading is enough, without
an advisory. That is backwards: judging how urgently to upgrade is what an
advisory is for.
Classification now decides whether there is an advisory. What operators
have to do decides what the advisory says.
---
security/index.html | 10 ++++++++--
1 file changed, 8 insertions(+), 2 deletions(-)
diff --git a/security/index.html b/security/index.html
index 080dcd8..f57ca6c 100644
--- a/security/index.html
+++ b/security/index.html
@@ -17,9 +17,9 @@ title: Reporting Security Problems to Apache StormCrawler
<h3>Crawled Content Is Untrusted by Design</h3>
<p>StormCrawler is a library for broad web crawling. It fetches URLs it
was told to fetch, including URLs discovered in previously fetched pages, and
it parses the bytes that come back. Both the URLs and the bytes are chosen by
the operators of the sites being crawled, which means attacker-controlled input
is not an edge case for a crawler: it is the normal operating condition. Every
component that handles a response body, a response header, a redirect target or
an extracted outlink is [...]
- <p>It follows that consequences of the crawler crawling what it was
configured to crawl are not, in themselves, vulnerabilities. If a topology is
pointed at a host, fetches a resource from it and stores the result, that is
the software working as intended, regardless of what the remote host chose to
serve. What the project does treat as a security problem is content crossing a
boundary it was never meant to cross: crawled bytes influencing the crawler's
own configuration or control data [...]
+ <p>It follows that consequences of the crawler crawling what it was
configured to crawl are not, in themselves, vulnerabilities. If a topology is
pointed at a host, fetches a resource from it and stores the result, that is
the software working as intended, regardless of what the remote host chose to
serve. Content crossing a boundary it was never meant to cross is what the
project fixes first: crawled bytes influencing the crawler's own configuration
or control data, reaching resources [...]
<p>This is the frame for the rest of this section. The controls
described below exist so that operators can define where that boundary lies for
their deployment. Choosing not to configure them widens the boundary rather
than removing it.</p>
- <p>Malformed or hostile responses that make a crawl slow, stall a
worker or exhaust its memory are bugs the project fixes. Such effects are
confined to the crawl the operator chose to run, and the project treats them as
robustness rather than as a compromise of the deployment.</p>
+ <p>Malformed or hostile input that makes a crawl slow, stalls a worker
or exhausts its memory is a bug the project fixes, whether it arrives in a
response or by way of data the crawl stored earlier. Availability effects of
this kind are confined to the crawl the operator chose to run, and the project
treats them as robustness rather than as a compromise of the deployment. This
is separate from the availability of the Apache Storm® cluster itself,
which is the operator's to protect.</p>
<h3>Trusted Configuration</h3>
<p>The configuration file used by StormCrawler is loaded during
topology submission and is treated as a trusted source. It does not involve any
user-supplied input at runtime.</p>
@@ -79,6 +79,12 @@ title: Reporting Security Problems to Apache StormCrawler
<li>We treat improvements to default configuration values and
added safeguards as security hardening. These ship in normal releases,
described in the release notes, without an advisory.</li>
<li>Vulnerabilities in dependencies follow the ASF's dependency
guidance. We update affected dependencies in normal releases and do not usually
issue our own advisory for them.</li>
</ul>
+ <p>Because StormCrawler is a library rather than a deployed service,
the configuration it ships is a starting point and not a secured deployment.
Where a weakness is reachable because a control described on this page was not
configured, we improve the default and treat the change as hardening. Where the
library defeats a control an operator did configure, or discloses a configured
secret in a way this page does not describe, we treat it as a vulnerability.</p>
+
+ <p>Whether we publish an advisory follows from that classification and
not from how much work the fix requires of anyone. A vulnerability in a
released artifact gets an advisory even when installing the next release is all
an operator has to do, because the point of an advisory is to let operators
judge how urgently to upgrade. Where more than an upgrade is needed, such as
changing a configuration, inspecting or repairing stored data, or rotating a
credential, the advisory says so and t [...]
+
+ <p>Log output is not a confidentiality boundary. Enabling DEBUG logging
on the protocol implementations can put configured credentials into the logs.
Treat DEBUG output from a topology that uses an authenticated proxy or crawls
authenticated sites as sensitive, and scope log retention accordingly.</p>
+
<p>As the ASF policy notes, what counts as a vulnerability depends in
part on what a project states as expected behaviour, which is why this page
describes the crawler's threat model rather than only its reporting process.</p>
<h3>Summary</h3>