This is an automated email from the ASF dual-hosted git repository.
rzo1 pushed a commit to branch security-model-expansion
in repository https://gitbox.apache.org/repos/asf/stormcrawler-site.git
The following commit(s) were added to refs/heads/security-model-expansion by
this push:
new 8aff87c Present the security model section as a settled proposal
8aff87c is described below
commit 8aff87cdceccaad4e10325a2315d4b9e4acb378c
Author: Richard Zowalla <[email protected]>
AuthorDate: Tue Aug 18 19:48:41 2026 +0200
Present the security model section as a settled proposal
Removes the draft status note and the three 'under discussion' markers so
the section reads as a coherent proposal for the PMC to accept, amend or
reject as a whole. The substantive operator guidance in those paragraphs
is kept: availability effects under hostile content are treated as
robustness, and cookie scoping is not a security boundary.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
---
security/index.html | 9 +++------
1 file changed, 3 insertions(+), 6 deletions(-)
diff --git a/security/index.html b/security/index.html
index 3995dc9..d320e05 100644
--- a/security/index.html
+++ b/security/index.html
@@ -15,13 +15,11 @@ title: Reporting Security Problems to Apache StormCrawler
<h2>Threat Model and Security Considerations</h2>
<p>StormCrawler is designed to operate in trusted environments as part
of a distributed Apache Storm® cluster. This document outlines the threat
model and key security assumptions to help users understand the secure use and
deployment of StormCrawler.</p>
- <p><strong>Status:</strong> this section is a draft under discussion on
the StormCrawler PMC's mailing lists. The paragraphs marked <em>Under
discussion</em> record positions that have been proposed but not yet settled,
and should not be read as project policy until this note is removed.</p>
-
<h3>Crawled Content Is Untrusted by Design</h3>
<p>StormCrawler is a library for broad web crawling. It fetches URLs it
was told to fetch, including URLs discovered in previously fetched pages, and
it parses the bytes that come back. Both the URLs and the bytes are chosen by
the operators of the sites being crawled, which means attacker-controlled input
is not an edge case for a crawler: it is the normal operating condition. Every
component that handles a response body, a response header, a redirect target or
an extracted outlink is [...]
<p>It follows that consequences of the crawler crawling what it was
configured to crawl are not, in themselves, vulnerabilities. If a topology is
pointed at a host, fetches a resource from it and stores the result, that is
the software working as intended, regardless of what the remote host chose to
serve. What the project does treat as a security problem is content crossing a
boundary it was never meant to cross: crawled bytes influencing the crawler's
own configuration or control data [...]
<p>This is the frame for the rest of this section. The controls
described below exist so that operators can define where that boundary lies for
their deployment. Choosing not to configure them widens the boundary rather
than removing it.</p>
- <p><em>Under discussion:</em> whether the crawler's own availability
under hostile content is a security guarantee or a robustness property.
Malformed or hostile responses that make a crawl slow, stall a worker or
exhaust its memory are unambiguously bugs the project fixes, but the PMC has
not yet decided whether such self-inflicted availability effects are treated as
security issues, and this page will be updated when it has.</p>
+ <p>Malformed or hostile responses that make a crawl slow, stall a
worker or exhaust its memory are bugs the project fixes. Such effects are
confined to the crawl the operator chose to run, and the project treats them as
robustness rather than as a compromise of the deployment.</p>
<h3>Trusted Configuration</h3>
<p>The configuration file used by StormCrawler is loaded during
topology submission and is treated as a trusted source. It does not involve any
user-supplied input at runtime.</p>
@@ -70,10 +68,9 @@ title: Reporting Security Problems to Apache StormCrawler
<h3>Transport and Credentials</h3>
<p><code>http.trust.everything</code> currently defaults to
<code>true</code>. With that setting, the OkHttp protocol accepts any server
certificate and performs no hostname verification. The rationale is that broad
crawling encounters expired, self-signed and otherwise misconfigured
certificates constantly, that a crawl aborting on them would be of limited use,
and that the fetched bytes are treated as hostile in any case, so certificate
validity adds little to how the response is hand [...]
<p>That rationale covers the content of an ordinary anonymous crawl. It
does not cover anything the crawler sends. With certificate and hostname
verification disabled, the crawler cannot tell which host it is actually
talking to, so any credential, custom header or replayed cookie attached to a
request crosses a channel whose far end was never authenticated. Operators who
crawl anything authenticated must weigh this, and should set
<code>http.trust.everything</code> to <code>false</code [...]
- <p><em>Under discussion:</em> whether permissive TLS is a posture the
project declares and keeps, with the setting made more visible in shipped
configurations and at runtime, or a default that should simply be changed. Both
options are on the table; the PMC has not yet chosen between them. Operators
should not assume the current default is permanent.</p>
<p>Credentials configured through <code>http.basicauth.user</code> and
<code>http.basicauth.password</code> are turned into a fixed
<code>Authorization</code> header that is attached to every request the
protocol instance makes, with no host scoping. If such a topology follows an
outlink to a third-party host, the credentials go with it. Operators crawling
authenticated sites should scope credentials by origin themselves, by running a
separate topology per origin with URL filters confin [...]
<p>Cookie replay is a related risk. <code>http.use.cookies</code> is
<code>false</code> by default, and enabling it causes cookies stored from a
response to be sent with requests for links found in that response. Operators
enabling it on an authenticated crawl should apply the same per-origin
separation as for Basic authentication.</p>
- <p><em>Under discussion:</em> whether the cookie scoping described in
the project's internals documentation constitutes a commitment about which
cookies are sent where. The PMC has not yet decided how to treat it. Until it
does, operators should not rely on cookie scoping as a security boundary and
should confine authenticated crawls by topology instead.</p>
+ <p>Cookie scoping as implemented is not a security boundary. Operators
should confine authenticated crawls by topology rather than relying on which
cookies are sent where.</p>
<p>Finally, proxy credentials configured via
<code>http.proxy.user</code> and <code>http.proxy.pass</code> can appear in log
output when the protocol implementations log at DEBUG level. DEBUG logging on a
topology configured with an authenticated proxy should be treated as sensitive,
and log retention scoped accordingly.</p>
<h3>Advisories and Hardening</h3>
@@ -83,7 +80,7 @@ title: Reporting Security Problems to Apache StormCrawler
<li>We treat improvements to default configuration values and
added safeguards as security hardening. These ship in normal releases,
described in the release notes, without an advisory.</li>
<li>Vulnerabilities in dependencies follow the ASF's dependency
guidance. We update affected dependencies in normal releases and do not usually
issue our own advisory for them.</li>
</ul>
- <p>As the ASF policy notes, what counts as a vulnerability depends in
part on what a project states as expected behaviour, which is why this page
describes the crawler's threat model rather than only its reporting process.
Where behaviour described here is still under discussion, the classification of
a report about that behaviour may be as well.</p>
+ <p>As the ASF policy notes, what counts as a vulnerability depends in
part on what a project states as expected behaviour, which is why this page
describes the crawler's threat model rather than only its reporting process.</p>
<h3>Summary</h3>
<p>StormCrawler's security model assumes a trusted deployment
environment. Users should:</p>