Elen Le Foll 🇫🇷 🇬🇧 🇩🇪
Postdoc and lecturer in corpus linguistics and project leader in the CRC 1252 "Prominence in Language" at the University of Cologne (Germany)
Personal account. All posts CC BY-NC.
#CorpusLinguistics #AppliedLinguistics #EFL #ELT #SLA #Rstats #OpenScience
• Former research fellow at the Centre for English Corpus Linguistics at UCLouvain
• Ph.D. in English linguistics/English language teaching
• Conference Interpreter M.A.
• Cognitive Science M.Sc.
• she/her •
📚📊🎷🎼🌱🍰🚴♀️🚄
I was only made aware of this (frankly awesome) case of LLM poisoning today: https://www.nature.com/articles/d41586-026-01100-y. A researcher made up a disease and published two evidently fake preprints about it (including sentences such as “this entire paper is made up” and “Fifty made-up individuals aged between 20 and 50 years were recruited for the exposure group”), which were almost immediately picked up by LLMs and documented in their output. Worse, actual – supposedly serious – medical papers also started citing the preprints, demonstrating that academics relying on LLMs to do their work is a genuine problem! Not that I had my doubts but, if anyone did, this seems like the perfect demonstration of the problem. Article immediately added to the syllabus of the class I am co-teaching with Iris Ferrazzo on LLMs for Romance Studies/Humanities!
#LLM #GenAI #academia #research #ResearchIntegrity #humanities
It was an early morning start for me as a volunteer at the Music Summer School Festival in Norfolk as I was on breakfast duty. Now in the choir rehearsal: wish us luck for the Gloria from Bach’s Mass in F minor!
mssf.org.uk
You are cordially invited to take the Rorschach tree test.
I suggest everyone put a content warning on their answers to make it more exciting!
I just submitted my textbook manuscript to @langscipress@openbiblio.social! 🎉🎉🎉 I published the initial version of the first few chapters online in April 2024 and this project has been at the centre of my attention for much of these past two years so it's a big day for me! I am incredibly grateful to the many students and colleagues whose critical and encouraging feedback has been instrumental to improving the book. The online version is here: https://elenlefoll.github.io/RstatsTextbook/.
My #WindowFriday today is a flashback to last year’s Music Summer School and Festival at Gresham‘s in Holt, North Norfolk, England. I can’t wait for this year’s summer school where I will, for the 10+ year, help out as a houseparent volunteer. If you’re an amateur musician or singer, of any age, background, or musical ability, check out the programme and come and join us! https://www.mssf.org.uk/
Germany has a problem: most outlets seem to sell Reibekuchen in packs of three. Personal research has shown that three Reibekuchen is definitely too much for one person. At the same time, 1.5 Reibekuchen is clearly not enough for a meal. Sign my petition to ban the sale of odd numbers of Reibekuchen!
I have really enjoyed following #DHd2026 on Mastodon – a huge thank you to everyone who tooted! Next semester, #ReproducibiliTea in the HumaniTeas (https://ub.uni-koeln.de/en/courses-consultations/reproducibilitea-in-the-humaniteas) continues in Cologne and online. We would love to see more #DH participants from DACH and beyond!
Join our mailing list to be the first informed: https://lists.uni-koeln.de/mailman/listinfo/reproducibilitea-humaniteas
If you'd like to contribute a session (a 20-min talk or 90-min workshop), get in touch ASAP as we are currently planning next semester's schedule! 🫖
This great toot reminded me of a conversation I overheard a few months ago between an acquaintance who works for an insurance company and one who works at a hospital in a clinical profession: one thought that AI was great for writing letters to companies citing laws to impress them, but rubbish (and even potentially dangerous) when it came to giving medical advice, while the other thought the exact opposite. No points for guessing who thought what!
RE: @ElenLeFoll@fediscience.org
This is today, my friends! Be there or be square! We are live-streaming, but not recording to ensure that everyone can speak their mind, share their anecdotes, and generally feel free to discuss even what may, at first, seem like a silly idea!
Yesterday and today I cycled a few more kilometres than usual to attend the Fourth International Conference on Prominence in Language. I enjoyed some great discussions with colleagues and international guests and some pretty decent views on the way there and back!
I have received news from the publisher of my 2024 book on Textbook English (J. Benjamins) that a Chinese university press has expressed interest in publishing a Chinese translation. I think this is very exciting, but also quite scary because I will have absolutely no way of assessing the quality of the translation. The publishers are currently negotiating and will inform me of the terms later on, but is there anything I should watch out for?
Last week, I had the pleasure of participating in a small conference on corpus linguistics and AI organised by colleagues from English and German linguistics at the University of Gießen. I am grateful to my co-author @mshakir_Dr@mastodon.world for suggesting that we present our work there as I wouldn't have applied otherwise and to the organisers for allowing different perspectives and opinions to be heard. The format allowed for lots of time to discuss our diverging views in a stunning setting.
Leute, das mit der Steuersenkung für die Inlandsflüge, das ist doch ein Aprilscherz, oder?!?!
Any bright ideas out there? 💡 Here's an "essay competition inviting active scientists from any sector to share concrete research challenges that you hypothesize are caused by broad structural or systemic bottlenecks in science, and experimental strategies to fix them": https://astera.org/essay-competition/. The first prize is 30,000$!
EDIT: Please read @ced@mapstodon.space's reply below. It turns out "active scientists" don't include economists, sociologists, psychologists, historians, and presumably anyone else doing social sciences and humanities research! 🤬
While preparing for the next session of my LLM class on training data, I came across this brilliantly illustrated article from the @washingtonpost@mstdn.social analysing the content of Google’s C4 data set, a filtered version of the Common Crawl used as training data for many LLMs: https://www.washingtonpost.com/technology/interactive/2023/ai-chatbot-learning/ (free access). Added to the seminar's required reading list!
This is now the second time I notice that, having paid by card at a self-checkout in #Lidl, I am charged twice for two different amounts. Both times I bought some snacks for meetings and sadly did not keep the receipts, so I don't know if they add up to what I bought or not. Does anyone have any idea why this might be happening?