From what I know of HDT, it supports one access pattern very well
(linked data fragments), and it is good for published large datasets on
the web. If there is enough memory to store some of the access
structures it would be OK for SPARQL. Without all indexes., as the
RAM/disk ratio gets worse so will SPARQL performance if the queries are
not of the right shape.
Note that hdt-jena is LGPL3 licensed.
In terms of sheer speed, parsing RDF Binary is 800K triples/s (for me),
N-Tripes is 200 kTPS.
> a 30M triple NT file. Loading that into TDB would probably take hours.
So about 3M triples? Or ii sthat comrpressed size? (my rule of thumb is
x8-x10 for N-triples).
The loading rate for TDB tdbloader should be about 30-50 kTPS (the rate
drops with increasing size). Faster for smaller datasets - 50-70K triples/s.
TDB loading is doing work "up front" so that any access patter in a
query is well served.
The easiest way to speed it up would be to remove some of the indexes.
TDB internally can cope with any combination of indexes, if there is at
least one (the primary).
If you are getting much slower "something is wrong". Running on a
laptop with a rotating disk will be slower, especially if you are using
the machine at the same time. Laptop SSDs aren't always that great
(it's a cost thing about how they interface to the system).
It is faster to work with RDF Binary - pure parsing is 800K triples/s
and it is fast to write. Files are large, like N-triples, it compresses
(gzip) well (x8-x10).
gzip compression gets some of the benefits of HDT - it finds common
symbols and has a dictionary - but it does not have path access. It is
universally available.
Andy
TBD2 is faster to load (not radically), especially loading into an
existing dataset.
On 04/04/17 10:56, Osma Suominen wrote:
Hi,
I have some experience using HDT with Jena. I think HDT is an amazing
technology and I've so far been happy with the performance, but as Rob
said, the use case matters a lot and benchmarking is recommended.
In my case I have a conversion pipeline [1] that converts a set of MARC
bibliographic records into a 30M triple NT file. Loading that into TDB
would probably take hours. Instead I'm converting it to HDT using the
hdt-cpp toolkit [2] (it's faster than the Java version and uses less
memory) and create an index file alongside the main HDT file. The
HDT+index files are a fraction of the size of the NT file (4GB NT file
vs. less than 500MB for the HDT+index).
I can then start up a version of Fuseki that exposes the data in the HDT
file as a read-only SPARQL endpoint. In my experience, query performance
is very reasonable, though I haven't benchmarked it against TDB. Since
the HDT file + index are rather small, they will soon be held mostly in
the disk cache, so although the technology is disk based, in practice
the disk will not be used very much unless you are extremely low on
memory.
Running the conversion from NT to HDT, creating the index file, and
starting up Fuseki altogether take less than 5 minutes and SPARQL
queries can then be run immediately. In that time the TDB loader would
have barely started.
As an alternative to Fuseki, SPARQL queries can be run directly on the
HDT file using the hdtsparql command line tool from the hdt-jena toolkit.
-Osma
[1] https://github.com/NatLibFi/bib-rdf-pipeline
[2] https://github.com/rdfhdt/hdt-cpp
04.04.2017, 12:27, Rob Vesse kirjoitti:
HDT is primarily on disk. Whether it is query-able depends on the
exact encoding, there is one encoding designed primarily for
transportation of data and another designed for querying called
HDT-FoQ aka focused on querying
In either case, there will be some memory usage as they do perform
some caching. They may also take advantage of memory mapped files
similar to what TDB does.
As far as comparisons with TDB I have never done any myself. For
simplistic queries, I would expect that HDT performs ok since from
what I remember the indexing is suitable for simple scans. However,
for queries with any kind of complexity i.e. Filters, negations, Joins
etc. I would expect TDB to outperform it and will scale far better.
But as I always point out on these kinds of questions your use case
will matter. If you think one solution will be better than the other
for your use case then you should benchmark that yourself. Generic
benchmarking will only tell you so much and give you a general
indication of comparative performance.
Rob
On 04/04/2017 07:03, "Lorenz B." <[email protected]>
wrote:
Well, I'm not that familiar with HDT, thus, I'm probably wrong. And I
saw right now that they also provide some kind of indexing concept.
Let's wait for response from Andy and/or Rob.
(In the meantime, I'll play around with HDT and Jena today to get
some
more insights. )
>> Jena HDT is in-memory, right?
> Is it? I thought it was a on-disk, compressed, and query-able
list of quads...
>
--
Lorenz Bühmann
AKSW group, University of Leipzig
Group: http://aksw.org - semantic web research center