https://bugs.koha-community.org/bugzilla3/show_bug.cgi?id=22972

--- Comment #51 from David Cook <[email protected]> ---
(In reply to [email protected] from comment #50)
> $0 indeed works for the local authority link. But $1 in bibliographic
> records is still relatively rarely used in production, despite being defined
> for Real World Object URIs. The fact that it "works semantically" per the
> MARC documentation for a different purpose doesn't mean this use is wrong -
> it just means it hasn't been exercised this way yet. Every extension of
> practice has to start somewhere.

I'm confused... isn't the point of these patches to add $1 subfields?

> > "Except if we look at the Library of Congress [...] they don't bring in all 
> > those valuable external identifiers into the Bib/Work. They keep them in 
> > the Authority record which is linked from the Bib/Work."
> 
> That's true for the LoC ecosystem, where every consumer of bibliographic
> data is assumed to also be able to resolve the authority side via
> id.loc.gov. That's not the reality for many Koha installations. Our bib data
> constantly leaves the system without the authority records attached: OAI-PMH
> harvesting, MARCXML export to union catalogues, indexing into
> Solr/Elasticsearch for discovery layers, data exchange with APIs, etc. All
> of those consumers see only the bib record. If the external ID lives solely
> in the authority, it's invisible to that entire chain unless every
> downstream consumer sets up a live link into our Koha authority table -
> which in practice nobody does.

But you were saying before that you want to create Linked Open Data. If you're
going to use a local authority, then it would make sense to link to that local
authority. That's why I think we should shift from the $9 for an authid to a $0
with a canonical URL to the authority record. 

> > "To me, if you wanted those 024$a identifiers in the $1, I don't see why 
> > you wouldn't copy them by hand. [...] I don't see the merit of including 
> > them automatically."
> 
> Copying by hand undermines exactly what authority control is for: one point
> of maintenance, automatically consistent everywhere the authority is used.
> With hundreds or thousands of bib records linked to a single authority,
> manual upkeep isn't realistic, and any correction or merge of the authority
> would again require touching every linked record by hand. That's precisely
> the problem authority control already solves for $a/$9; we're only asking to
> extend that same logic to $1.

What I mean is... you don't need all those $1s but if you want to add them you
can add them. Why would a correction require touching every linked record by
hand? You've already made the link via the URL. Unless you're talking about
changes to URLs. Then you're already in a different realm of bulk/batch
changes.

> > "Essentially, copying the Auth 024$a into the $1 of the Bib record is just 
> > denormalising the data, and for what gain?"
> 
> Yes, it's denormalisation - just like the 6XX$a we already populate with the
> authority's preferred term, or the $9 we already populate with the authid.
> Denormalisation for the sake of portability is already standard practice in
> library data, not an exception. The gain is that the external identifier can
> leave the bib record without requiring an extra join.

No, the $9 is not denormalization. It's a linkage. It's the key to create that
"join". 

You do have me thinking again though... LoD can be a pain because following
links on-demand/at run-time is a real big pain. Much more efficient to
denormalize. That's true... but in a different context. It's true in the
context where you might periodically cache fetched data from external URLs to
improve performance. Especially if you're using a graph database.

In this Koha context, all you're doing is displaying the URLs. There's other
ways of easily fetching that data rather than embedding it into the bib record.

That said, I take your point. You want an easy way to automate the population
of $1 URLs in a bib record from the 024$a URLs in the authority record. While I
don't see the merit overall, I can understand the merit in terms of your
specific goal for sure.

But it seems to me that a plugin to the cataloguing editor would make more
sense here, because really that's what the change is all about, right? You want
to show the data via the XSLTs and you want to copy the URLs from the authority
record to the bib record at catalogue time. 

I suppose what I'm trying to say is... this change adds to the existing
technical debt by adding more code on an already poorly designed mechanism.
While it's a small change to the existing system, it's adding to the burden. We
want to be easing the burden. But easing the burden takes more work in the
short-term so it's rarely done.

> > "If it's 3rd party tools, why would they be interested in the other $1 
> > URLs? Surely, it would be enough to link to 1 $1 URL to create the linkage."
> 
> There isn't one universal "best" identifier: one consumer wants Wikidata (as
> we do – it's a citizen science-compliant recognized Digital Public Good),
> another wants VIAF (which is mostly included in Wikidata entities), another
> wants ISNI, etc. Multiple $1s let each consumer pick up the identifier
> relevant to their use case, without us as a library having to decide in
> advance which consumer we're serving.

I think part of the problem here is that you're thinking in terms of
identifiers instead of concepts. 

There is no universal "best" identifier, but in practice for libraries you
typically would link to an authority record. Traditionally, that authority
record has lived in Koha, but it would be interesting to link to external
authorities instead. (Now that would be a really great new feature to help
support Linked Open Data in Koha cataloguing!)

What are the use cases of your consumers? Real or hypothetical; I'm curious. As
a library, what you're doing is you're saying "Hey this is a resource we have,
and this metadata description I'm using has an authoritative form which you can
find in X URL location". 

> Finally, you're right that this sits in core and that the configuration is
> at the authority-type level rather than a system preference. That seems like
> a fair point to carry into further QA - an explicit system preference to
> toggle this globally could address the concern about "silent, bespoke
> behaviour."

I think I would be more amenable to this change if there were a system
preference for it - for sure. 

> But the core point stands for us: without this, external identifiers stay
> locked in the authority silo, invisible to anything outside Koha itself.
> That's exactly the problem the Semantic Web is meant to solve, and it's why
> we consider this worth pursuing further - not out of stubbornness, but
> because it answers a real interoperability problem.

I disagree. 

For instance, take this real life document that uses LOD from the University of
Vienna using Fedora Commons: https://phaidra.univie.ac.at/detail/o:2353677 and
https://phaidra.univie.ac.at/api/object/o:2353677/json-ld

Specifically, look at this section:
  "edm:hasType": [
    {
      "@type": "skos:Concept",
      "skos:exactMatch": [
        "https://pid.phaidra.org/vocabulary/PYRE-RAWJ";
      ],
      "skos:prefLabel": [
        {
          "@language": "eng",
          "@value": "other"
        }
      ]
    }
  ],

So the LoD URL is https://pid.phaidra.org/vocabulary/PYRE-RAWJ. That is the
canonical identifier for an "authority record". If you follow that link, you
will see that at that canonical URL there are "Links to other vocabularies",
which includes http://purl.org/coar/resource_type/c_1843 and then that includes
the concept rendered in many different languages.

Now in terms of denormalisation, they've included the "eng" label here. So they
can use the English label easily without needing to dereference the LoD links. 

Note on the HTML UI you do not see any links to
http://purl.org/coar/resource_type/c_1843. All you see is that English label
"other" linking to https://pid.phaidra.org/vocabulary/PYRE-RAWJ

Does that make sense? That's how Linked Open Data systems work in the real
world. 

So when you have a $0 https://mykoha/authority/1 it links that bib record to
that authority record. Someone (or something if it's a bot) can then go to that
authority record and go "Oh it's equivalent to this VIAF record or this
WikiData record". 

That's what the Semantic Web is all about. Or rather having machine-readable
machine-understandable linkages using canonical URLs to link objects with
semantic meaning is what the Semantic Web is all about. 

The idea is to provide URLs to metadata in structured formats that machines can
understand and navigate to connect related content in automated ways. 

> Happy to keep working toward a form that's also acceptable to you - a system
> preference layered on top, for instance.

For what it's worth, I'm not trying to be a gatekeeper. If Jonathan wanted to
pass QA and Pedro as release manager wanted to push it to main, I could not and
would not stop them. 

Rather, I'm trying to bring all my knowledge and experience in cataloguing and
linked open data to this discussion.

-- 
You are receiving this mail because:
You are watching all bug changes.
_______________________________________________
Koha-bugs mailing list -- [email protected]
To unsubscribe send an email to [email protected]
website : http://www.koha-community.org/
git : http://git.koha-community.org/
bugs : http://bugs.koha-community.org/

Reply via email to