Hi Jakob,

I would like to chime in here because I recently faced a similar
requirement.  I wanted to extract and reorganize some existing MarkLogic
content to create on-demand reports.  As I looked into doing this with a
CPF, I found that I wanted to be able to update and reopen XML documents
while looping through the various documents containing the source
information.  I ran into problems because of the nature of transactions.

I talked with one of the professional support staff who asked what I was
trying to accomplish.  I was trying to create relationships between data in
an element and in an attribute in order to get counts, and to provide a
drill down into the specific documents containing both sets of data.  He
pointed me to the cts:element-attribute-value-co-occurrences() function
which creates a list of co-occurrences of the element and attribute, each of
which needed to be configured in a lexicon (for which I use the range
element and attribute indexes) and for which I had to set the fragment root
appropriately (as the lexicon search functions work off fragments, not
documents).  Accordingly I was able to use the cts:frequency() function to
get counts from the lexicon and from the results of the co-occurrences.  As
a result I am able to cook up web reports using a MarkLogic HTTP server in
pretty short order without having to extract and store information for
analysis purposes elsewhere.

For the drill down I was able to pull the paired results form the
co-occurrences output and to perform a cts:search to find the XML documents
that match the information - pretty cool beans!

This approach may not be suitable for your needs, but I thought I'd throw it
into the mix for your consideration.

Regards,

Tim Meagher - AAOM Consulting

-----Original Message-----
From: [email protected]
[mailto:[email protected]] On Behalf Of Jakob Fix
Sent: Thursday, June 25, 2009 9:04 AM
To: General Mark Logic Developer Discussion
Subject: Re: [MarkLogic Dev General] force a "buffer write"?

Thanks David and Geert,

I'm currently looking into xmlsh. David, would there be sample scripts
available somewhere which could be used by me?  Geert, no I haven't
looked at neither triggers nor cpf, but I will explore this route too.

thanks for your input.
Jakob.



On Thu, Jun 25, 2009 at 14:54, Geert Josten<[email protected]> wrote:
> Hi,
>
> I agree that it does not necessarily make sense to do this within
MarkLogic. On the other hand, you might have a good reason. Particularly
when there is really lots of information to analyse, and you have to store
your information somewhere..
>
> Have you considered taking an asynchronized approach? You can have one
query gather all the uri's that need processing and store that somewhere,
perhaps in batches. Then use Triggers or CPF to process those batches.
MarkLogic is capable of handling those triggers in multiple threads, though
if all uri's point to the same website, you perhaps don't want to overload
it that way. Perhaps a small sleep would be appreciated by the website
hoster..
>
> You could also use scripts and programming languages to do stuff, but then
it might be better to do that part outside MarkLogic all together, and only
insert the log reports for analysis purposes..
>
> Kind regards,
> Geert
>
>> -----Original Message-----
>> From: [email protected]
>> [mailto:[email protected]] On Behalf Of
>> Lee, David
>> Sent: donderdag 25 juni 2009 13:50
>> To: General Mark Logic Developer Discussion
>> Subject: RE: [MarkLogic Dev General] force a "buffer write"?
>>
>> I've done link-checking programs before and I suggest this
>> may be best done *outside* of ML.
>> What I would do if I were to do this
>>
>> 1) Use ML to generate an XML document with the info you need
>> (xml file with list nodes)
>> 2) Use a scripting language, or programming language that
>> supports multithreads or multiprocessors
>> 3) In batches of N threads/processes test the links and write
>> the results to non-conflicting output files
>> 4) wait for each batch to continue and aggregate the results
>> (maybe send this bit back to ML ?)
>> 5) Goto 3 until done
>> 6) Send the aggregated results back to ML
>>
>>
>> For the scripting or programming language for #2, there are
>> many options .
>> My personal bias, of course, is xmlsh which runs background
>> tasks as threads, but sh, perl, java , C++ or any language
>> that lets you do background processing of URL fetches will work.
>> One that will use threads instead of processes is desirable
>> but I've done this kind of thing with sh before and its
>> acceptable.  You may want a URL fetch command that has a
>> settable timeout, something like wget, many urls' that are
>> inaccessible may take 30 seconds or more to time-out using
>> default timeouts.
>> OTOH if you use a language with effecient trheading you can
>> do large batches (say 100 or more in parallel) and the
>> individual timeouts wont matter as much.
>>
>> -David Lee
>>
>>
>>
>>
>>
>> -----Original Message-----
>> From: [email protected]
>> [mailto:[email protected]] On Behalf Of
>> Jakob Fix
>> Sent: Thursday, June 25, 2009 7:39 AM
>> To: General Mark Logic Developer Discussion
>> Subject: Re: [MarkLogic Dev General] force a "buffer write"?
>>
>> Thanks once more, Geert!
>>
>> >From an architectural point of view, does it make sense to
>> have a loop
>> over many thousand URLs run in an xquery which may take 5
>> seconds each?  How does Mark Logic handle long-running
>> queries?  The goal is to assemble the information about the
>> accessibility of these URLs in an XML document that will be
>> stored in Mark Logic and used for analytical output.  Should
>> the list be fragmented into smaller sub-lists and be
>> processed separately?
>>
>> Would it be possible to have several threads run simultaneously?
>> Somehow I doubt it as there would probably be issues with the
>> final aggregation of the different thread sub-documents into
>> a big one.
>>
>>
>> What about timeouts, especially if the function is called
>> from inside a web page, how does Mark Logic handle this issue
>> (I saw that one can tweak the timeout), or would this be a
>> browser timeout problem?
>>
>> I hope these questions are not too off-topic for the list.
>>
>> Jakob.
>>
>>
>> On Thu, Jun 25, 2009 at 08:11, Geert Josten
>> <[email protected]> wrote:
>> >
>> > Hi Jakob,
>> >
>> > You are looking for unbuffered response streams, but
>> sending of the response is handled fully by the HTTP server.
>> I don't believe you can influence that.
>> >
>> > Giving it some more thought, I am afraid that allowing
>> unbuffered responses would break the idea of transactions.
>> You don't want to send back response, unless you can
>> guarantee no exceptions will be thrown. And I don't think
>> that can be guaranteed.
>> >
>> > Perhaps a MarkLogic expert would like to comment?
>> >
>> > Hasn't this been discussed before? It vaguely rings a bell..
>> >
>> > Kind regards,
>> > Geert
>> >
>> > >
>> >
>> >
>> > Drs. G.P.H. Josten
>> > Consultant
>> >
>> >
>> > http://www.daidalos.nl/
>> > Daidalos BV
>> > Source of Innovation
>> > Hoekeindsehof 1-4
>> > 2665 JZ Bleiswijk
>> > Tel.: +31 (0) 10 850 1200
>> > Fax: +31 (0) 10 850 1199
>> > http://www.daidalos.nl/
>> > KvK 27164984
>> > De informatie - verzonden in of met dit emailbericht - is
>> afkomstig van Daidalos BV en is uitsluitend bestemd voor de
>> geadresseerde. Indien u dit bericht onbedoeld hebt ontvangen,
>> verzoeken wij u het te verwijderen. Aan dit bericht kunnen
>> geen rechten worden ontleend.
>> >
>> >
>> > > From: [email protected]
>> > > [mailto:[email protected]] On
>> Behalf Of Jakob
>> > > Fix
>> > > Sent: donderdag 25 juni 2009 1:39
>> > > To: General Mark Logic Developer Discussion
>> > > Subject: [MarkLogic Dev General] force a "buffer write"?
>> > >
>> > > So, I've written a function that looks at one URL at a time and
>> > > returns true or false depending on its accessibility.
>> > > Now, my problem is that the result is returned only when all (and
>> > > that means potentially many) URLs have been checked.
>> > > Isn't there a way to "force a write"? I'm not sure I'm expressing
>> > > myself correctly, but hopefully you'll understand what I mean.
>> > > Thanks.
>> > >
>> > > (: consider this: if the timeout is 10 seconds, and I have three
>> > > "bad" URLs, I may have to wait up to 30 seconds before seeing the
>> > > results, isn't there a way to see a result every ten seconds
>> > > instead? :)
>> > >
>> > > for $url in $urls
>> > >     return <xh:li>{$url}:
>> > > {utils:http-resource-available($url, $timeout)}</xh:li>
>> > >
>> > >
>> > > (: function that checks accessibility of a URL, written for DOIs
>> > > which are "placeholder" URLs which forward to the real URL :)
>> > >
>> > > declare function utils:http-resource-available
>> > >     ($doi as xs:string, $oldtimeout as xs:integer?) as
>> xs:boolean {
>> > >     try {
>> > >
>> > >       let $timeout := if ($oldtimeout) then $oldtimeout else 10
>> > >
>> > >       let $head :=
>> > > xdmp:http-get(fn:concat($utils:doi-resolver, $doi),
>> > >           <options
>> > > xmlns="xdmp:http"><timeout>{$timeout}</timeout></options>)
>> > >       let $code := $head//xdh:code cast as xs:integer
>> > >       let $location := $head//xdh:location
>> > >       return ((fn:contains($location,
>> > > $utils:location-part-to-match)) and
>> > >           ($code < 400)) (: we want 3XX or 2XX? :)
>> > >     } catch ($ex) {
>> > >         if ($ex/error:code eq 'SVC-SOCRECV')
>> > >         then
>> > >             fn:false()
>> > >         else
>> > >             xdmp:rethrow()
>> > >     }
>> > > };
>> > >
>> > >
>> > >
>> >
>> > _______________________________________________
>> > General mailing list
>> > [email protected]
>> > http://xqzone.com/mailman/listinfo/general
>> _______________________________________________
>> General mailing list
>> [email protected]
>> http://xqzone.com/mailman/listinfo/general
>> _______________________________________________
>> General mailing list
>> [email protected]
>> http://xqzone.com/mailman/listinfo/general
>> _______________________________________________
> General mailing list
> [email protected]
> http://xqzone.com/mailman/listinfo/general
>
_______________________________________________
General mailing list
[email protected]
http://xqzone.com/mailman/listinfo/general


_______________________________________________
General mailing list
[email protected]
http://xqzone.com/mailman/listinfo/general

Reply via email to