Dan Goodman
mastodon 4.8.0-alpha.3+glitchI'm a computational neuroscientist and science reformer. I'm based at Imperial College London. I like to build things and organisations, including the Brian spiking neural network simulator, Neuromatch and the SNUFA spiking neural network community.
You might have read that arxiv is banning people for a year if they post LLM-generated papers and cheered it on. But most of the discussion about this doesn't correctly explain the policy and it is not a good thing.
First up, the policy is that "incontrovertible evidence" of using LLMs and not checking the output is what's at stake. An example given is a hallucinated reference.
Second, the ban will apply to all coauthors of the paper, not just the person submitting.
Third, it's not just a 1 year ban, it's followed by a permanent ban on submitting papers that have not been peer reviewed in a "reputable" journal or conference. Given that arxiv is a preprint server and not a repository for published papers, that makes the ban effectively permanent.
(EDIT: They've subsequently clarified that this isn't a permanent ban but will be lifted after 1 or 3 (the clarification isn't clear) peer reviewed papers. This is still problematic, but much better. The rest of this post left as I originally wrote it.)
So imagine: you are a masters student working on a project with a few other people, and your role is relatively minor. The project leads to a paper and you get your name on it, hurray. The lead author handles the submission and doesn't ask for your permission to send the final version because you're only a masters student. Your supervisor explains that this is how things are done and nothing to worry about. What you didn't know is that someone else on this paper at the last minute made some edits to the grammar of the paper using an LLM because none of you are native English speakers, and the LLM inserted a hallucinated reference. Arxiv picks up on this and you are now permanently banned from using arxiv as a preprint server. Further, every time you try to collaborate with someone else to write a paper and they want to put it in arxiv you have to explain that you can't, and that this means that by collaborating with you, they also can't put it on arxiv. Since arxiv is one of the main channels for distributing papers in your field, soon enough people stop asking you to work with them and your career is effectively over. Because someone else didn't notice that an overenthusiastic grammar checker inserted a fake reference and you weren't in a position of enough power at the time to insist on checking the final version.
Ok that's a long story, but I don't think this is a fanciful situation. Stuff like this happens all the time. It's easy to say - and I've seen a lot of people saying times like this - that everyone should take responsibility for reading the paper, or that it's the responsibility of supervisors to make sure this doesn't happen. But in the world we actually inhabit, power imbalances exist: the masters student can't make the supervisor wait until they read the paper because they're worried about their project grade. Bad supervisors are out there, and it's not fair to punish their students.
This policy will lead to terrible consequences for a lot of innocent people who should not reasonably be held responsible because they weren't in a position of power. I suspect it won't lead to very bad consequences for big name researchers who will just get on the phone to someone at arxiv or one of arxiv's funders and get the decision reversed in their case.
I understand the anger towards LLMs and tech companies, and I share it. I understand the anger towards the people cynically generating whole papers using them, polluting the scientific literature and making all our lives more difficult, and I share it too. But that doesn't mean we should jump to implement extreme and poorly thought out policies that will hurt a lot of people who haven't done anything wrong.
Finally, as an advocate of open science and publishing reform, this is really disappointing from arxiv. By saying that peer reviewed papers in "reputable" journals are ok, they've defined themselves (arxiv) as second class citizens in the world of publishing. This shows such limited ambition, and actively hurts the cause of making the world better by getting rid of the parasitic and harmful publishing industry.
@curvenote@hachyderm.io and @openrxiv@biologists.social @biorxivpreprint@biologists.social teaming up to make an incredibly cool new tool to read biorxiv preprints as HTML with nice features (particularly on desktop when you hover over links). Try it on any biorxiv preprint, for example my most recent one (with just a couple of bugs that I expect will get ironed out soon):
To me little quality of life features like this can potentially push the visibility and viability of preprints and non-journal infrastructure, so I'm quite excited about this.
I think there's some genuinely interesting conversations to be had about machine learning and mathematics. My problem is that every time I see an article on this topic now, it always has such strong branding - for example Tim Gowers' recent article has the product name and version in the title itself - and I can't help but feel like I'm the victim of an intricate (likely not direct) marketing scheme. That makes it hard for me to want to engage in that discussion because I have no interest in helping these companies sell their products.
Right, time for a social media break. Not sure how long. See you on the other side (or maybe not, who knows). Might drop back to post new papers.
RE: @thetransmitter@mastodon.social
Great article and I completely agree. If we want to study the interesting thing about the brain ("intelligence") we need to study it doing hard tasks, not easy ones that can be solved with a single neuron or even none. Yes it makes the analysis harder but that's the challenge.
Idea for grant proposal formative assessment. Very short proposal (1 page say) and reviewers are each allowed 3 questions which you get a very limited character count to respond to. Overall effort hugely reduced and lets the proposal evolve within a round rather than needing resubmission. Thoughts?
RE: @gconstantinides@mathstodon.xyz
We're hiring for a project on #SpikingNeuralNetworks and #neuromorphic computing, to start in October this year, for 36 months. Can hire at pre- or post-PhD level. Feel free to email me informally, or apply at the link below. Please do share with your networks if you know someone who would be interested.
#ComputationalNeuroscience
RE: @neuralreckoning@neuromatch.social
In general, academics talking about specifically named LLM products (rather than talking about LLMs in general) is now such a recognisable genre that it's hard to imagine that it's not somehow coordinated and planned as a marketing campaign. Would be fascinating to try to trace how this has happened.
RE: @thetransmitter@mastodon.social
The issues raised in this article bring to mind de Tocqueville's 1835 "Democracy in America" where he argues the reason for the success of American "democratic" science over European "aristocratic" science is that it makes things simple enough for anyone to reuse. We lose that using AI for science.
There may even be a similar social dynamic. Those with the wealth and power to make things more difficult for ordinary people have a strong incentive to do so, because it excludes those ordinary people from taking part and secures their dominant position.
@gepasi@fediscience.org sure, it's better that he's saying this than being gung ho on AI, but many people are saying this sort of thing, including many who are not personally profiting from exploitation of academic labour.
If you're interested in doing a 2 year Schmidt "AI in science" postdoc fellowship in neuroscience/AI stuff with me starting in July or Oct 2027 take a look at this and get in touch soon. We've had a lot of luck recently getting these fellowships.
https://www.imperial.ac.uk/electrical-engineering/research/fellowships/
RE: @thetransmitter@mastodon.social
This is a great article. I still feel less excited about "neural foundation models" than Juan, but I have to say he makes his case very convincingly, both for the advantages and disadvantages. Well worth a read!