https://bugzilla.wikimedia.org/show_bug.cgi?id=72507

            Bug ID: 72507
           Summary: Allow crawling of bugzilla.wikimedia.org select
                    content
           Product: Wikimedia
           Version: wmf-deployment
          Hardware: All
               URL: http://web.archive.org/save/https://bugzilla.wikimedia
                    .org/duplicates.cgi
                OS: All
            Status: NEW
          Severity: normal
          Priority: Unprioritized
         Component: Bugzilla
          Assignee: [email protected]
          Reporter: [email protected]
                CC: [email protected], [email protected],
                    [email protected], [email protected],
                    [email protected]
       Web browser: ---
   Mobile Platform: ---

The robots.txt rules are unnecessarily restrictive. As bugzilla is being
deprecated, and only a portion of its content migrated to phabricator, it's
essential that we allow third parties to do their job. All crawlers, or at
least ia_archiver (wayback machine), should be allowed to crawl:
1) any content which
2) doesn't specifically cause load issues and
3) is not being semantically migrated to phabricator.
Ideally we'd drop requirement (3) but let's start somewhere.

Example URLs which shouldn't be blacklisted:
* /page.cgi?id=voting/bug.html*
* /duplicates.cgi*
* /report.cgi* (unless load)
* /weekly-bug-summary.cgi*
* /describecomponents.cgi*

In fact, is there any reason not to allow everything, minus:
* /show_bug.cgi
* /showdependencytree.cgi
* /query.cgi
?

-- 
You are receiving this mail because:
You are the assignee for the bug.
You are on the CC list for the bug.
_______________________________________________
Wikibugs-l mailing list
[email protected]
https://lists.wikimedia.org/mailman/listinfo/wikibugs-l

Reply via email to