Boosted by @joe@f.duriansoftware.com
Replying to
Our core argument is that LLMs should not be regarded as reliable simulators of human participants because subtle changes to the input text can lead to very different reactions from humans and LLMs. This is also shown in our key figure below. The x axis corresponds to different moral scenarios, the y axis to the reaction if we subtly change the text of the scenario. You can see that the reactions of LLMs and humans are clearly very different for most scenarios (exact correlations in the paper).