https://bugzilla.wikimedia.org/show_bug.cgi?id=72507
Bug ID: 72507
Summary: Allow crawling of bugzilla.wikimedia.org select
content
Product: Wikimedia
Version: wmf-deployment
Hardware: All
URL: http://web.archive.org/save/https://bugzilla.wikimedia
.org/duplicates.cgi
OS: All
Status: NEW
Severity: normal
Priority: Unprioritized
Component: Bugzilla
Assignee: [email protected]
Reporter: [email protected]
CC: [email protected], [email protected],
[email protected], [email protected],
[email protected]
Web browser: ---
Mobile Platform: ---
The robots.txt rules are unnecessarily restrictive. As bugzilla is being
deprecated, and only a portion of its content migrated to phabricator, it's
essential that we allow third parties to do their job. All crawlers, or at
least ia_archiver (wayback machine), should be allowed to crawl:
1) any content which
2) doesn't specifically cause load issues and
3) is not being semantically migrated to phabricator.
Ideally we'd drop requirement (3) but let's start somewhere.
Example URLs which shouldn't be blacklisted:
* /page.cgi?id=voting/bug.html*
* /duplicates.cgi*
* /report.cgi* (unless load)
* /weekly-bug-summary.cgi*
* /describecomponents.cgi*
In fact, is there any reason not to allow everything, minus:
* /show_bug.cgi
* /showdependencytree.cgi
* /query.cgi
?
--
You are receiving this mail because:
You are the assignee for the bug.
You are on the CC list for the bug.
_______________________________________________
Wikibugs-l mailing list
[email protected]
https://lists.wikimedia.org/mailman/listinfo/wikibugs-l