Hi,

there were several articles by Peter Jasco in the past regarding search quality
in Google Scholar (GS), see for example

http://www.libraryjournal.com/article/CA6698580.html

The GS "inclusion guidelines" are about three years old now, so I'm wondering
about the discussion now. In the past, the number of documents from repositories
was even higher in Google than in GS! The experience with our own repositories
shows, that providing GS metadata clearly increased the number of documents of
our repositories in GS, but is still below 50%. BASE covers repository content
much better than GS, of course GS has other qualities (citation counts, library
links, ...).

>From my point of view the exciting question is, if GS uses the GS metadata only
to get the fulltext easier from a repository or if GS uses the metadata in
addition to the fulltext in order to improve the search quality within GS. The
other question is, how much the Google/GS ratio of documents from repositories
has changed in the last years.

Best
Dirk

------------------------------------------
Dirk Pieper
Bielefeld UL - BASE
Universitätsstr. 25, D-33615 Bielefeld
E-mail: [email protected] | Tel.: +49 521 106-4010
Fax: +49 521 106-4052

www.ub.uni-bielefeld.de
www.base-search.net
------------------------------------------


+++ Welcome to the 10th International Bielefeld Conference,
24. - 26. April 2012,
http://conference.ub.uni-bielefeld.de +++



----- Ursprüngliche Nachricht -----
Von: Stevan Harnad <[email protected]>
Datum: Samstag, 18. Februar 2012, 8:31
Betreff: [GOAL] {Disarmed} Re: Google Scholar discoverability of repository
content
An: "Global Open Access List (Successor of AmSci)" <[email protected]>
Cc: SPARC IR <[email protected]>

> Begin forwarded message:
>
      > From: Betsy Coles <[email protected]>
> Date: February 17, 2012 5:48:42 PM EST
> To: [email protected]
> Subject: Re: [EP-tech] Re: Google Scholar discoverability of repository
content
>
> I'm the technical manager for the main IR at Caltech, CaltechAUTHORS

      > (MailScanner has detected a possible fraud attempt from
      "authors.library.caltech.edu" claiming to be
      http://authors.library.caltech..edu), currently running EPrints
      3.1.3.  
      >
      > Tim's conjecture 1) below seems to account almost exactly for the
      result

      > the article authors found: 87.7% of the 25,072 eprints in
      CaltechAUTHORS

      > have OA documents attached; the remainder have only documents that
      are

      > either restricted to campus or to repository staff.  I don't think
      there are very

      > many cases of Tim's conjecture 2), since we have concentrated on
      adding

      > current content.
      >
      > I haven't read the article in question (we don't subscribe), but
      the percentage

      > of open access eprints is almost exactly the same as the authors'
      report of GS

      > indexed items in Table 2.  I haven't tested specifically, but it's
      tempting to

      > conclude that GS is indexing 100% of our open access content.
      >
      > Betsy Coles
      > Caltech Library IT Group
      > [email protected]
      >
      > -----Original Message-----
      > From: [email protected]
      [mailto:[email protected]] On Behalf Of Tim Brody
      > Sent: Friday, February 17, 2012 3:33 AM
      > To: [email protected]
      > Cc: [email protected]
      > Subject: [EP-tech] Re: Google Scholar discoverability of
      repository content
      >
      > Hi All,
      >
      > Here is some specific advice for existing repository
      administrators from Google Scholar:
      > http://roar.eprints.org/help/google_scholar.html
      >
      > As far as I'm aware there isn't anyone running EPrints 2 now, so
      EPrints-based repositories are already (and for a long) the "best in
      class" for Google Scholar.
      >
      >
      > Right, this paper ...
      >
      > Table 1 is irrelevant and misleading. Scholar links first to the
      publisher and, only if there is no publisher link, directly to the
      IR version. That's a policy decision on the part of Scholar and
      nothing to do with IRs.
      >
      > Table 2 gives us some useful data. The headline rate for EPrints
      is 88% (based on CalTech). Unfortunately the authors haven't
      provided an analysis of what happened to the missing records. I've
      done a quick random sample of CalTech and I suspect the missing
      records will consist
      > of:
      > 1) Non-OA/non-full-text records (I'm sure a query to the CalTech
      repository admin could supply the data).
      > 2) A percentage of PDFs that Scholar won't be able to parse.
      CalTech contains some old (1950s), scanned PDFs from Journals. Where
      the article isn't at the top of the page Scholar will struggle to
      parse the title/authors/abstract and therefore won't be able to
      match it to their records e.g.
      http://authors.library.caltech.edu/5815/
      >
      >
      > The remainder of the paper describes the authors' process of
      fixing their own IR (based on CONTENTdm).
      >
      >
      > The authors then wrongly conclude:
      >
      > "Despite GS’s endorsement of three software packages, the surveys
      conducted for this paper demonstrates that software is not a
      deciding factor for indexing ratio in GS. Each of the three
      recommended software packages showed good indexing ratios for some
      repositories and poor ratios for others."
      >
      > The authors looked at one instance of EPrints and, despite being a
      relatively old version, found 88% of its records indexed in GS.
      >
      > It is unfortunate that this paper has suggested that IR software
      in general is poorly indexed in GS. On the contrary, some badly
      implemented IR software is poorly indexed in GS.
      >
      >
      > After all that is said, the most critical factor to IR visibility
      is having (BOAI definition) open access content. Hiding content
      behind search forms, click-throughs and other things that emphasise
      the IR at the expense of the content will hurt your visibility.
      >
      > Lastly, Google will index your metadata-only records while Google
      Scholar is looking for full-texts. Your GS/Google ratio will
      approximate how many of your records have an attached open access
      PDF (.doc etc).
      >
      >
      > Sincerely,
      > Tim Brody
      > (EPrints Developer)
      >
      > On Wed, 2012-02-15 at 11:31 +0000, Stevan Harnad wrote:
            > Can we enhance the google-scholar discoverability of
            EPrints (and

            > DSpace) repositories?

            >

            >
            
http://linksource.ebsco.com/linking.aspx?sid=google&auinit=K&aulast=Ar

            >
            
litsch&atitle=Invisible+Institutional+Repositories:+Addressing+the+Low

            >
            
+Indexing+Ratios+of+IRs+in+Google+Scholar&title=Library+Hi+Tech&volume

            > =30&issue=1&date=2012&spage=4&issn=0737-8831

            >

            > Kenning Arlitsch, Patrick Shawn OBrien, (2012)
            "Invisible

            > Institutional

            > Repositories: Addressing the Low Indexing Ratios of
            IRs in Google

            > Scholar", Library Hi Tech, Vol. 30 Iss: 1

            >

            > Purpose - Google Scholar has difficulty indexing the
            contents of

            > institutional repositories, and the authors
            hypothesize the reason is

            > that most repositories use Dublin Core, which cannot
            express

            > bibliographic citation information adequately for
            academic papers.

            > Google Scholar makes specific recommendations for
            repositories,

            > including the use of publishing industry metadata
            schemas over Dublin

            > Core. This paper tests a theory that transforming
            metadata schemas in

            > institutional repositories will lead to increased
            indexing by Google

            > Scholar.

            >

            > Design/methodology/approach - The authors conducted
            two surveys of

            > institutional and disciplinary repositories across the
            United States,

            > using different methodologies. They also conducted
            three pilot

            > projects that transformed the metadata of a subset of
            papers from

            > USpace, the University of Utah's institutional
            repository, and

            > examined the results of Google Scholar's explicit
            harvests.

            >

            > Findings - Repositories that use GS recommended
            metadata schemas and

            > express them in HTML meta tags experienced
            significantly higher

            > indexing ratios. The ease with which search engine
            crawlers can

            > navigate a repository also seems to affect indexing
            ratio. The second

            > and third metadata transformation pilot projects at
            Utah were

            > successful, ultimately achieving an indexing ratio of
            greater than 90%.

            > Research limitations/implications - The second survey
            was limited to

            > forty titles from each of seven repositories, for a
            total of 280 titles.

            > A larger survey that covers more repositories may be
            useful.

            >

            > Practical implications - Institutional repositories
            are achieving

            > significant mass, and the rate of author citations
            from those

            > repositories may affect university rankings. Lack of
            visibility in

            > Google Scholar, however, will limit the ability of IRs
            to play a more

            > significant role in those citation rates.

            > Originality/value - Little or no research has been
            published about

            > improving the indexing ratio of institutional
            repositories in Google

            > Scholar. The authors believe that they are the first
            to address the

            > possibility of transforming IR metadata to improve
            indexing ratios in

            > Google Scholar.

            > *** Options:

            >
            http://mailman.ecs.soton.ac.uk/mailman/listinfo/eprints-tech

            > *** Archive: http://www.eprints.org/tech.php/

            > *** EPrints community wiki: http://wiki.eprints.org/

      >
      >
      > *** Options:
      http://mailman.ecs.soton.ac.uk/mailman/listinfo/eprints-tech
      > *** Archive: http://www.eprints.org/tech.php/
      > *** EPrints community wiki: http://wiki.eprints.org/

>
> _______________________________________________
> GOAL mailing list
> [email protected]
> http://mailman.ecs.soton.ac.uk/mailman/listinfo/goal



    [ Part 2: "Attached Text" ]

_______________________________________________
GOAL mailing list
[email protected]
http://mailman.ecs.soton.ac.uk/mailman/listinfo/goal

Reply via email to