Skip to content

Instantly share code, notes, and snippets.

@alasdairham
Created July 21, 2017 05:07
Show Gist options
  • Select an option

  • Save alasdairham/76401e2dd3bc1606553a4bed669ff328 to your computer and use it in GitHub Desktop.

Select an option

Save alasdairham/76401e2dd3bc1606553a4bed669ff328 to your computer and use it in GitHub Desktop.
class Player:
# ....
def update_strategy(self):
"""
Set the preference (strategy) of choosing an action to be proportional to positive regrets. e.g, a strategy that prefers PAPER can be [0.2, 0.6, 0.2]
"""
self.strategy = np.copy(self.regret_sum)
self.strategy[self.strategy < 0] = 0 # reset negative regrets to zero
summation = sum(self.strategy)
if summation > 0:
# normalise
self.strategy /= summation
else:
# uniform distribution to reduce exploitability
self.strategy = np.repeat(1 / RPS.n_actions, RPS.n_actions)
self.strategy_sum += self.strategy
def learn_avg_strategy(self):
# averaged strategy converges to Nash Equilibrium
summation = sum(self.strategy_sum)
if summation > 0:
self.avg_strategy = self.strategy_sum / summation
else:
self.avg_strategy = np.repeat(1/RPS.n_actions, RPS.n_actions)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment