new paper: "ToxTempAssistant: using large language models to standardise cell-based toxicological test method descriptions" https://doi.org/10.1080/2833373X.2026.2638036
"This study quantifies the tool’s baseline performance under controlled conditions. ToxTempAssistant uses grounded, per-question prompting with mandatory source attribution. Evaluation paired a positive control (..) with a negative control (..) across three LLM models (gpt-4.1-nano, gpt-4o-mini, o3-mini)."