The GitHub Actions job "Java CI with Maven" on 
stormcrawler.git/issue-2087-robots-content-limit has failed.
Run started by GitHub user abhinav-phi (triggered by rzo1).

Head commit for run:
6f2730ba7e2ad7ae22ec075f4ff751a4cedba889 / abhinav-phi <[email protected]>
Ignore invalid metadata content limits, reword the widening rationale (#2087)

Review feedback on #2121:

- a http.content.limit in the metadata below -1 is a configuration
  error and is ignored instead of being applied as a nonsensical limit
- the comment on the -1 special case now says what the shipped default
  of http.robots.content.limit actually does: the robots.txt fetch is
  deliberately allowed a larger read than pages when
  http.content.limit is below the 512 kiB RFC floor, because the floor
  is the reason the robots specific key exists
- confirmed no other protocol reads the http.content.limit metadata
  key: only the okhttp HttpProtocol does, WARCSpout reads the config
  value directly and is unaffected

Report URL: https://github.com/apache/stormcrawler/actions/runs/34153040205

With regards,
GitHub Actions via GitBox

Reply via email to