Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I don't think that follows. To beat the machine the move must be both unpredicted and profitable. Random moves are not profitable. Training purely by reinforcement learning rather than on humans could create a policy network that ignores more subtrees that are profitable than the current one does. In short, it isn't good enough for the AI to be good at playing itself, it has to be good at playing every possible player, and while it is playing humans it is sufficient for it to be good at playing every human player.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: