ainews.cx Entities

OpenAI News

Measuring Goodhart’s law

Goodhart’s law famously says: “When a measure becomes a target, it ceases to be a good measure.” Although originally from economics, it’s something we have to grapple with at OpenAI when figuring out how to optimize objectives that are difficult or costly...

Topics: Models

Entities: OpenAIModelsTarget

OpenAI News

Equivalence between policy gradients and soft Q-learning

Two of the leading approaches for model-free reinforcement learning are policy gradient methods and Q-learning methods. Q-learning methods can be effective and sample-efficient when they work, however, it is not well-understood why they work, since...

Topics: Policy

Entities: PolicyTarget