The problem with scraping is not so much scraping one page but that
people do lots of pages fairly quickly and it impacts the service.
Joshua
Gijs Kruitbosch wrote:
Lindsay Donaghe wrote:
Yes, the API only lets you retrieve your own posts. That's to
prevent a lot of potential bandwidth abuse from what I understand...
It was a bummer when I realized that because I wanted to write a tool
that would compare my archive with my friends and point out links
that they had that I didn't that I might be interested in...
Oh well...
Be careful with screen scraping as an alternative because that's not
looked on highly either and could get your account banned.
Lindsay
So, this raises another few questions, that is, if your understanding of
the matter is correct :-)
* Why is bandwidth an issue? Surely 15 posts from myself or 15 posts
from 'the world' take up the same amount of bandwidth? I would assume
that any such API had the same or more restrictions on it than the
user-only one, to prevent abuse which would result in high server
CPU/Memory usage (that is, I'm assuming it might take more time/cpu to
crawl through *all* the posts than it takes to just get your own.
Maybe.). I'd even be fine with having a maximum of 500 or maybe even 100
queries to it a day, on one user account, and of course the same kind of
maximum that the current API has (default 15, and no more than 100 posts
per reply, IIRC?). You might want to throttle-enforce 5-second delays
between this and the user-only API, instead of 1-second. I'm quite fine
with that, and I think most people who would put it to fair use would
be. :-)
* Why is screen scraping (assuming that's the right word to use for
parsing the html manually and extracting data from it) not looked on
highly either? You only use this more-bandwidth-consuming way of doing
things because there is no suitable API... If there was, things would be
better, right?
* How would it get anyone banned? Whoever uses it would use it on
his/her own account, and the only thing that could possibly be done is
ban the user-agent. Which would be unfair, since all you're doing is
making a normal HTTP request, just as if you'd browse there normally
with a web browser (in the case of getting the html and parsing that,
that is).
* Why would requesting these posts using the API cause more server
load/bandwidth-loss than just using the search field on the webpage?
Forgive me if I seem rude or selfish here, I'm just not fully
understanding the issue that's said to be the cause of this lack in
the API.
Finally, it seems a lot of places (eg. digglicious, and anything
mentioned on here
(http://pchere.blogspot.com/2005/02/absolutely-delicious-complete-tool.html))
somehow do manage to get all the public bookmarks. I'm thinking they're
probably screenscraping too, or am I wrong there?
Again, I'm just curious and still hopeful I'll be able to use
del.icio.us in my app. If anyone can shed some more light on the
questions I asked, then please do.
Best regards,
Gijs Kruitbosch
PS: Argh, got to get used to posting to mailing lists again. Sorry for
sending this to you privately too, Lindsay.
_______________________________________________
discuss mailing list
[email protected]
http://lists.del.icio.us/cgi-bin/mailman/listinfo/discuss
_______________________________________________
discuss mailing list
[email protected]
http://lists.del.icio.us/cgi-bin/mailman/listinfo/discuss