Hi Remi,

If I understand correctly, your method makes your program 250 Elo points
stronger than my pattern-learning algorithm on 5x5 and 6x6, by just
learning better weights.

Yes, although this is just in a very simple MC setting.

Also we did not compare directly to the algorithm you used (minorisation/maximisation), but to a different algorithm for maximising the likelihood of an expert policy. But as the softmax policy is a log transformation of the generalised Bradley-Terry model, I think this comparison is still valid.

I had stopped working on my program for a very long time, but maybe I will go back to work and try
your algorithm.

I would be really interested to know how this works out! Please let me know.

Best wishes
-Dave

_______________________________________________
computer-go mailing list
[email protected]
http://www.computer-go.org/mailman/listinfo/computer-go/

Reply via email to