> *blinks* when I view the page it still has a
> tagcloud of common tags in the
> top-right...
> 

The data is not complete and the historical data is
missing. I shall explain more later.

> > Joshua once wrote - admittedly in a bit of a
> different context (viz. in
> > explaining that the "/post url allows only a title
> and url ... [in order] to
> > prevent XSS hacks," i.e., lest "publishers" use
> the feature to "suggest"
> > tags):
> >
> > Tags are a memory aid, decided upon by the person
> doing the remembering.
> > > If some other person with intentions other than
> recall choses the tags, then
> > > they are indeed improperly chosen.
> > >
> >
> > If tags are "memory aids" - and memory aids only -
> why bother displaying
> > *all* of the tags from each user?... The "Common
> tags" ought to be good
> > enough to get the mnemonic juices flowing.


Tags as just 'memory aids' is a very superficial point
of view. It might have been true when Joshua wrote it,
but now no one believes it anymore. Delicious was a
succesfull experiment, and when it started no one
realised the importance that the tags would have
reached. It would have been like imagining the taste
of a fruit by seeing a seed of the plant for the first
time.

Joshua might have said that they were just memory
aids, but later he said that people pay their acces to
the site by tagging things. Why would he be interested
in having those tags if they were only 'memory aids'.

On the 25 of May 2005 I showed how mathemathically
your 'memory aids' could be used to compute a distance
between web pages. 
http://blog.pietrosperoni.it/2005/05/25/tag-clouds-metric/

Four month before 'Ben Hyde' had proved something that
many people had been suggesting from a long time. That
is that tags followed a power law. He actually
calculated the steepness of the power law.
http://enthusiasm.cozy.org/archives/2005/01/tagging-powerlaw/

Of course he did this by screen scraping, because this
kind of info was not at the time given through the API
(and still isn't). Much research have been done in the
meantime to study folksonomy, and in general the
behaviour of your 'visual aids'. This research was
generally done either directly screen scraping, or
using websites that were screen scraping. The whole
geek community got its fruit, as new web sites would
appear using tags. Joshua, didn't all this helped you
at least a bit in understanding the depth, and the
importance of tags, and thag clouds? Well this
research was only possible BECAUSE delicious have been
offering the whole historical data of the tag cloud.
By screen scraping you could not only see the tag
cloud of now, but what was the tag cloud in every
moment. And people who were screen scraping, yes,
against the guidelines, but always telling to Joshua,
and following the guidelines were giving back to the
community. 

We would use delicious, delicious would gather the
data, the data would then be offered in html format
and researchers would use the data to study the tags,
tag clouds, and folksonomy and would feed back in the
community their discoveries.

Three days later from the previous post of tag clouds
metric, Terrell Russell (who now shared his
frustration with me via email) made a web site:
Cloudalicious. Cloudalicious could show how the tag
cloud would change as the days would pass.

By studying this I could MEASURE HOW SOCIETY WAS
CHANGING. The dream of many sociologist, and we had it
available on a golden plate. I presented my work the
morning after in another entry that also became very
popular: 

http://blog.pietrosperoni.it/2005/05/28/tagclouds-and-cultural-changes/

In this I showed how you can see the changes in the
society as they are happening, and indeed measure
them. I took as example the tag ajax.

All this was only possible because cloudalicious was
offering the data. And cloudalicious was offering the
data because delicious was making it available.
Delicious stops it, and a very promising line of
academic research dies.

The line got his first academic paper (that I know of)
from Golder and Huberman, on the 18th of August. They
even studied the tag clouds of every url that came out
from delicious popular in a period of 4 days. 

http://arxiv.org/abs/cs.DL/0508082

They made some very respectable work. Reproduced some
previous results and proposed reasons why tag clouds
would follow a power law. This was of course only
possible because delicious were permitting people to
have the whole tag cloud for each single URL, and
indeed the whole historical data of the tag cloud.

You might ask yourself why are people so keen on using
delicious respect to, let's say Flickr. It is not only
that delicious is bigger than many other system, it is
also that the information was more readily available
and more cleanly kept. In delicious every user had a
tag set of each url, and each url ended up having a
tag cloud made up of all the tag sets. I will not
enter here on why this is important, please refer to
my first blog entry above to see the details. For now
believe me, it is fundamental.

The fact that people used delicious in their example
all the time meant a huge free advertisment for
delicious. An advertisment that would generally reach
a different community from the geek community from
which delicious was born.

And an advertisment that delicious does not seem to
value much anymore.

Many many people used this information. Saying that
only few people were using it totally forgets that
many people were using cloudalicious, and that needed
that information. And many people were reading the
results from the people making the research thus
indirecly having a benefit from all the data being
offered back to the community. It is like saying that
since only 3% of personal computer run linux than
linux is unimportant. Totally forgetting that 99% of
the people uses google, and google runs on linux. 

This line of academic research is not dead... yet.

3 weeks ago I was giving a presentation at the
university on bottom-up onthologies. I contacted Clay
Shirky for the occasion. In that occasion he told me:

"I've been giving a talk pointing to your 'little
social quake' graphic and trying to convince everyone
that this is the most important view of tags currently
going."

My advisor looking at the results asked if I shouldn't
aim to make some serious research for a Nature
publication or similar. If you are in academic you
know that for many discipline Nature is the top to
which many people aspire. Above: the nobel prize. He
actually asked if I shouldn't write a research
proposal for the EU once the PhD was over. 

People writing academic papers; studies being done;
people inspired by those studies starting new web
sites; research; ideas; the possibility to calculate
the distance between url; the possibility to
investigate the long tail of the web through that; the
possibility to study social changes through studying
tag clouds. Delicious getting free advertisment from
all those studies. Advertisment reaching way beyond
the geek community it started (many professors at my
talk never heard of delicious before).

And you are being serious in saying that this data is
unnecessary to the majority of people. Well, until the
majority of people don't enjoy the fruit of the
research being done with it.


I asked Joshua 6 times now to make an API that
delivers that information. He NEVER answered. Not yes,
not no. Just ignored the message. The last was the one
from yesterday evening. It got ignored, again.

> PS - Last thought: There's probably a Greasemonkey
> solution to this
> > problem, eh?... Or not?... Could you not
> reconstruct the old page using
> > GM?...

Sure by calling the page of each user who bookmarked a
site. For the big site (which are the ones
statistically relevant) this means calling a thousand
pages for each url... plus the fact that for many user
you will need to go back in time through some pages.
It would be less abusive for delicious if we were to
start spidering the whole site and just make a copy of
the whole database.

As a last element I want to point out that the idea of
taking of the info because they wer cluttering the web
pahe, and not really being useful is ungrounded. There
are many information that we pick up from our
environment subconsciously. The underline emotional
tone of an email. (Could you feel that I was pissed
off yesterday? Can you feel I am more calm, but more
lcear today?). The expression of a face. This is
analog information. It uses a higher bandwidth to
digital information, but if you are able to use it it
gives you way more information than the digital one.
It is like a mouse respect to a keyboard. An ipod
wheel respect to using 4 buttons to move through
directories.

The info that delicious provided was giving a lot of
subconscious info to the reader. Look at the tags of a
user. How often do you read all the tags? Never,
right. Would you say that they give you no information
about that user. Can't you immediatly see what are his
interests, his way of thinking, his favorite topics? 

And so is the info about the URL, That info should not
have been taken away, it should have been made
complete, instead. 

And I totally subscribe that the name of the people
tagging something gives no info. You can as well just
put a number.

Ok, it took me 90 minutes to write this email.
If you read it all, I thank you.
If you think I am write please add your voice. If you
think I am not please tell me why too.

And let's hope that it is really a decision that was
coming out of Joshua, and he was not just obeying some
directives from the Yahoo headquarters.

Pietro



__________________________________________________
Do You Yahoo!?
Tired of spam?  Yahoo! Mail has the best spam protection around 
http://mail.yahoo.com 
_______________________________________________
discuss mailing list
[email protected]
http://lists.del.icio.us/cgi-bin/mailman/listinfo/discuss

Reply via email to