Ask a chatbot something dangerous and it will refuse the request. But this critical tool of AI safety is far from foolproof and could become an instrument of repression.