Hi Remi,
If I understand correctly, your method makes your program 250 Elo
points
stronger than my pattern-learning algorithm on 5x5 and 6x6, by just
learning better weights.
Yes, although this is just in a very simple MC setting.
Also we did not compare directly to the algorithm you used
(minorisation/maximisation), but to a different algorithm for
maximising the likelihood of an expert policy. But as the softmax
policy is a log transformation of the generalised Bradley-Terry model,
I think this comparison is still valid.
I had stopped working on my program for a very long time, but maybe
I will go back to work and try
your algorithm.
I would be really interested to know how this works out! Please let me
know.
Best wishes
-Dave
_______________________________________________
computer-go mailing list
[email protected]
http://www.computer-go.org/mailman/listinfo/computer-go/