Hi Doug,
do you know apache mina?
http://directory.apache.org/subprojects/network/
this is a nice introducing.
http://directory.apache.org/subprojects/network/mina.pdf

I know you are trying to stay as much as possible independent from third party libraries, but may
but mina looks very interesting from my point of view.

Stefan

Am 10.11.2005 um 21:03 schrieb Doug Cutting:

I've recently had the opportunity to experiment with Nutch on a 200
machine cluster running linux 2.4 kernels. It works well, but a larger
cluster might have problems using a 2.4 kernel.

The NDFS master (namenode) has a thread per connection.  Every task,
tasktracker and datanode process keeps a connection open to the
namenode.  2.4 linux kernels are limited to around 800 threads per
process.  With 200 boxes this means we can run two tasks per box at a
time (2 tasks, 1 tasktracker, and 1 datanode = 4 connections per box).
With 300 boxes, we might not even be able to run one task per box, since
that would result in 900 connections (1 task, 1 tasktracker, and 1
datanode per box).

Last week I looked into rewriting org.apache.nutch.ipc.Server to use a
single nio-based listener thread and a fixed number of worker threads.
I got it working, but it was slower and considerably more complex. The
complexity is because Nutch's object i/o is based on blocking,
stream-based i/o.

If I were to work more on it, here's how I might do it:

  . associate a pipe with each connection;
  . keep a queue of connections with new requests
  . keep a set of connections with requests in progress

  . listener loops, selecting connections for input
    for all connections with new input
      if the connection is in requests in progress set
        read the input and write to its pipe, potentially blocking
        (but not for long, since someone is reading it)
      else
        remove the connection from the selector
        queue the connection

  . worker threads loop
      pop a connection off the queue
      add it to the selector
      add it to the requests in progress set
      loop
        read a request from the pipe (potentially blocking)
        after request, if no more input is available in pipe
          remove it from the requests in progress set
          break
        compute the response
        write the response to the connection

This may not always be fair.  If a given client sends requests without
pause then there is the possibility that this client can starve other
clients.  In practice I don't think this would be a problem.  But I
don't see how to avoid it, since socket read boundaries may not
correspond to request boundaries.

I'm not sure this is worth working more on any more, since 2.6 kernels
can easily handle 10,000 or more threads.

Doug




-------------------------------------------------------
SF.Net email is sponsored by:
Tame your development challenges with Apache's Geronimo App Server. Download
it for free - -and be entered to win a 42" plasma tv or your very own
Sony(tm)PSP.  Click here to play: http://sourceforge.net/geronimo.php
_______________________________________________
Nutch-developers mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/nutch-developers

Reply via email to