As a researcher with many articles in the ACM library, I have to say this is a masterclass in hypocrisy. Obviously, lawyers can decipher the terms of ACM publishing contracts and Creative Commons licences to determine if this will be acceptable or not. But ACM is not a company, it's a non-profit founded in 1947 to represent scientists.
I would be surprised if a majority of ACM members were to say yes should we ask them (but ACM is not known for such democracy). Along with book authors, we are one of the many people that provide the knowledge and expertise on which large tech firms train their models, and get nothing in return. Actually, life is getting worse for us: extra workload in universities with students' AI use, a completely broken peer review system, etc. Hence the irony of ACM thinking about licensing, and only licensing, at a time where this is the least of our priorities.
The issue here is ACM focusing on licensing. A non profit would not be able to pay ACM for access. Hence why this policy is hypocritical: it gives more power to the larger players and undermines smaller actors in the field who have fewer resources.
Why did you sign off your copyright to ACM then? If you kept the copyright and prevented others from distributing your work, you could be the sole distributor and negotiate the price with the AI company yourself.
As long as we are going toward a world of abundance where money doesn't mean much and the main currency is time, I can't complain. I will subsidize that with my brain power turned into ink on paper.
Well, once there is no more jobs, or once a large enough number of people are made redundant, it will mechanically happen.
Future is all about research and entertainment. That is why tiktok stars and soccer players get so much money. All that being based on DARPA research, I mean the internet. People just want to have fun. Du pain et des jeux.
Abundance for everyone or it is not abundance. With abundance, the field gets levelled because everything someone else can get, you can too.
Society can't be equal when everyone is in survival of the fittest mode because resources (energy) are not infinite.
If you don't hold a patent for the use of the knowledge you published publicly, you can't prevent others from using the knowledge. You enjoy the prestige attached to the idea that you're an academic who participates in giving away their knowledge but then you play this game when that knowledge would actually be useful as opposed to being read by 3 other people in your special area who sit on your various committees in your career, now you want to forbid the use for culture war intra-elite signaling reasons.
You don't own the knowledge you put out there unless you have a limited time valid patent. The rest is absurdity. If you want to keep your findings to yourself, keep them secret.
The intellectual property law that governs ACM articles is copyright law, not patents. I don’t know who controls these (the ACM or the authors) or what rights might have been granted to the public.
The entire point of copyright law is so that people can make their writing public and still be able to control the right to make copies (for example, into your dataset for training an LLM).
The verbatim reproduction is clearly a red herring and not the main use case. Nobody reads novels (or science papers) by prompting ChatGPT to give the next paragraph.
Derivative work or transformative? It's not the same.
AI code generation often outputs exact copies of code that exists in the wild. I've seen it output chunks from research papers unprompted as well, its a big problem, or blending two papers together in a salad
AI works also clearly aren't transformative in many cases. If you ask it a question about a paper, it'll quote bits of the paper at you. That serves as an exact substitute of the original work. If you ask it for song lyrics, or information about the news, its content is a direct substitute for the original source it was trained on. This clearly does not fall under a transformative use case
You could argue that some uses of it are transformative, but even then - its easy to find some piece of training data in the source code that the output work supersedes. By its very nature it does not have the capacity to genuinely invent under the law (as it is not human), and a prompt isn't a significant enough part of the processing to count here
AI has the capacity to exactly reproduce its training data, just because something is a transformed representation does not mean that it isn't copying it in some fashion. The JPEG format 'just' counts the frequencies in an 8x8 block of pixels, and yes that's 100% copyright infringement
This makes no sense whatsoever. How would a philosopher of science, or a social scientist who publish in the ACM apply for a patent? Or someone who builds software (software patent not so easy to get ;)). I have applied for patents before and I'm pretty sure my patent application has been fed to countless LLMs by now.
The issue is not who owns knowledge, it's how it benefits humanity.
They can't apply for a patent in those categories, so they have no way of preventing others from reading their text (or using a program to process the text, and compute statistical properties), and using the ideas in other contexts (without re-expressing the same text).
If someone reads a philosophical essay, has a heureka moment from it, applies the principle to their work, and makes bank (commercial profit), they never have to pay a percentage to the author of the essay.
The world would be quite different if AI companies had to create the knowledge they trained on, rather than consume that knowledge freely given away. They're not known for freely giving away their produce either, I don't know why you think the ire should be pointing in this direction.
I like open models for sure. I support free software as well. But using published knowledge to solve new problems was never disallowed, even for profit. Today, a for-profit company, e.g. a gigantic Big Pharma company can have their employees read chemistry and biology papers and use the knowledge gained from it to improve their products and processes and make more profit without paying a cent to the authors (beyond what they may get - likely nothing - due to the potential paywall).
“If you don’t lock your bike, you can’t prevent others from taking it for a ride. You enjoy the mobility attached to the idea you’re a bike rider who rides a bike but then you play this game when that bike would actually be useful as opposed to sitting in the bike rack all day”.
It seems that the "open access" is more marketing than something reality based they actually want to do.
https://dl.acm.org/openaccess ---> So how to access the content? Do I have to register or what? It the "open access" only for academic org's people or for everyone in the world?
https://dl.acm.org/ ---> Okay, nice simple search field without loggin in, but when you try to search something, you get thousands of results, even if you search specific author and the exact name of paper, you will get hundreds of results and the thing you want is buried somewhere on page 247. Filters of authors, years etc. are for Premium subscription. But just googling the thing finds the link to ACM... And want to get the actual PDF? Hope it says "free access"...
Google scholar has become the way to search papers (which is somewhat worrisome). What people want from there is a download for all the stuff that does not list an author copy. Open access IMHO is just a reaction to the fact that mostly you would not need a subscription anyhow. Now the authors are paying upfront or universities are paying flat for all their researchers. The problem is now the incentives are not increasing the number of subscriptions but increasing the number of papers published.
I'm really not sure how this would work. I don't know how the ACM works, but in IEEE you would have to give them your publishing rights. However, training a LLM is not publishing by itself, it is a derivative work? Any way, at this point authors should be entitled to monetary compensation, not the publisher. The deal is totally different.
Of course they did. There is open access to this library. These guys actually think they're offering new training data? It's kind of hilariously naive.
ACM has been leaning heavily into AI-generated content for their journals in the past year, and this article is no exception: it appears to be 100% AI generated and full of LLM verbiage.
There's something hilarious about that, but also, snake eating its own tail.
Since starting to use AIs seriously for search in the last 3 months I have read and referenced more published papers and academic primary sources then I think I did in the previous 5 years. They're fantastic for pointing at some claim and asking for the primary source for it, and then it does the work of following things through the layers of backref to the original, assuming it's online. Opening the library up to AI access makes it far more accessible and usable then it was before.
AIs, at least in their current form, make you more who ever you were. If you want snap, glib answers of dubious accuracy, they'll give them to you, more easily than ever before. If you want to dig back into primary sources and get the original content, they'll do that for you, more easily than ever before.
Can't speak to how the science infrastructure is going to handle them, but if it takes down the peer review system, which I think has been worthless for probably going on two decades and has just given the entire enterprise a false sense of assurance, it'll probably be a net gain in the end. Peer review is a source of more problems then it is solving right now.
TLDR: “we’re going to try to get money from LLM providers for access to our back catalog without getting permission from the authors or providing them with any share of the revenue”.
I don't want to be mean but if it's that hard to even recognize that LLMs are even a valid thing that could intersect with your business that you need some kind of campaign for it..
It's almost like, "don't hurt yourself unc, we will just search arxiv".
As a researcher with many articles in the ACM library, I have to say this is a masterclass in hypocrisy. Obviously, lawyers can decipher the terms of ACM publishing contracts and Creative Commons licences to determine if this will be acceptable or not. But ACM is not a company, it's a non-profit founded in 1947 to represent scientists.
I would be surprised if a majority of ACM members were to say yes should we ask them (but ACM is not known for such democracy). Along with book authors, we are one of the many people that provide the knowledge and expertise on which large tech firms train their models, and get nothing in return. Actually, life is getting worse for us: extra workload in universities with students' AI use, a completely broken peer review system, etc. Hence the irony of ACM thinking about licensing, and only licensing, at a time where this is the least of our priorities.
If it was a non profit that trained the model - would that change your mind?
The issue here is ACM focusing on licensing. A non profit would not be able to pay ACM for access. Hence why this policy is hypocritical: it gives more power to the larger players and undermines smaller actors in the field who have fewer resources.
Why did you sign off your copyright to ACM then? If you kept the copyright and prevented others from distributing your work, you could be the sole distributor and negotiate the price with the AI company yourself.
You mean the knowledge you gathered with public grants, with a public paid salary, yet don’t want to make freely available to the public?
Yeah, too bad
As long as we are going toward a world of abundance where money doesn't mean much and the main currency is time, I can't complain. I will subsidize that with my brain power turned into ink on paper.
It’s pretty clear that the premise won’t be becoming reality.
Well, once there is no more jobs, or once a large enough number of people are made redundant, it will mechanically happen. Future is all about research and entertainment. That is why tiktok stars and soccer players get so much money. All that being based on DARPA research, I mean the internet. People just want to have fun. Du pain et des jeux.
Abundance for whom, exactly? And if the answer is everybody: who in charge has an incentive to do this?
Sorry to ruin your day, but if the people with money could have introduce equal society, you would have noticed their attempts by now.
Abundance for everyone or it is not abundance. With abundance, the field gets levelled because everything someone else can get, you can too. Society can't be equal when everyone is in survival of the fittest mode because resources (energy) are not infinite.
well, we are not, so far it is just concentrating wealth and power in smaller and smaller subset of people
If you don't hold a patent for the use of the knowledge you published publicly, you can't prevent others from using the knowledge. You enjoy the prestige attached to the idea that you're an academic who participates in giving away their knowledge but then you play this game when that knowledge would actually be useful as opposed to being read by 3 other people in your special area who sit on your various committees in your career, now you want to forbid the use for culture war intra-elite signaling reasons.
You don't own the knowledge you put out there unless you have a limited time valid patent. The rest is absurdity. If you want to keep your findings to yourself, keep them secret.
The intellectual property law that governs ACM articles is copyright law, not patents. I don’t know who controls these (the ACM or the authors) or what rights might have been granted to the public.
The entire point of copyright law is so that people can make their writing public and still be able to control the right to make copies (for example, into your dataset for training an LLM).
Copyright protects against reprinting or reproducing the wording and expression, not using the idea expressed in there in novel contexts.
AI reproduces copyrighted work exactly in many cases, so it clearly infringes copyright in this sense
The output of it is also a derivative work, and derivative works also infringe copyright. Its only not a problem if you ignore copyright entirely
Humans are the only entities that get to enjoy special idea-learning-exemptions, not AI
The verbatim reproduction is clearly a red herring and not the main use case. Nobody reads novels (or science papers) by prompting ChatGPT to give the next paragraph.
Derivative work or transformative? It's not the same.
AI code generation often outputs exact copies of code that exists in the wild. I've seen it output chunks from research papers unprompted as well, its a big problem, or blending two papers together in a salad
AI works also clearly aren't transformative in many cases. If you ask it a question about a paper, it'll quote bits of the paper at you. That serves as an exact substitute of the original work. If you ask it for song lyrics, or information about the news, its content is a direct substitute for the original source it was trained on. This clearly does not fall under a transformative use case
You could argue that some uses of it are transformative, but even then - its easy to find some piece of training data in the source code that the output work supersedes. By its very nature it does not have the capacity to genuinely invent under the law (as it is not human), and a prompt isn't a significant enough part of the processing to count here
Suddenly it's "copyright infringement" to count the amount of times one word occurs after another word. I find this whole thing so amusing.
AI has the capacity to exactly reproduce its training data, just because something is a transformed representation does not mean that it isn't copying it in some fashion. The JPEG format 'just' counts the frequencies in an 8x8 block of pixels, and yes that's 100% copyright infringement
The courts don't agree with you and I don't either. Now what?
This makes no sense whatsoever. How would a philosopher of science, or a social scientist who publish in the ACM apply for a patent? Or someone who builds software (software patent not so easy to get ;)). I have applied for patents before and I'm pretty sure my patent application has been fed to countless LLMs by now.
The issue is not who owns knowledge, it's how it benefits humanity.
They can't apply for a patent in those categories, so they have no way of preventing others from reading their text (or using a program to process the text, and compute statistical properties), and using the ideas in other contexts (without re-expressing the same text).
If someone reads a philosophical essay, has a heureka moment from it, applies the principle to their work, and makes bank (commercial profit), they never have to pay a percentage to the author of the essay.
The world would be quite different if AI companies had to create the knowledge they trained on, rather than consume that knowledge freely given away. They're not known for freely giving away their produce either, I don't know why you think the ire should be pointing in this direction.
I like open models for sure. I support free software as well. But using published knowledge to solve new problems was never disallowed, even for profit. Today, a for-profit company, e.g. a gigantic Big Pharma company can have their employees read chemistry and biology papers and use the knowledge gained from it to improve their products and processes and make more profit without paying a cent to the authors (beyond what they may get - likely nothing - due to the potential paywall).
Patents shouldn’t exist.
“If you don’t lock your bike, you can’t prevent others from taking it for a ride. You enjoy the mobility attached to the idea you’re a bike rider who rides a bike but then you play this game when that bike would actually be useful as opposed to sitting in the bike rack all day”.
Nope, and I'd also download cars.
How about we give humans access
Do humans not have access? https://dl.acm.org/openaccess
It seems that the "open access" is more marketing than something reality based they actually want to do.
https://dl.acm.org/openaccess ---> So how to access the content? Do I have to register or what? It the "open access" only for academic org's people or for everyone in the world?
https://dl.acm.org/ ---> Okay, nice simple search field without loggin in, but when you try to search something, you get thousands of results, even if you search specific author and the exact name of paper, you will get hundreds of results and the thing you want is buried somewhere on page 247. Filters of authors, years etc. are for Premium subscription. But just googling the thing finds the link to ACM... And want to get the actual PDF? Hope it says "free access"...
Nobody uses their search anyways.
Google scholar has become the way to search papers (which is somewhat worrisome). What people want from there is a download for all the stuff that does not list an author copy. Open access IMHO is just a reaction to the fact that mostly you would not need a subscription anyhow. Now the authors are paying upfront or universities are paying flat for all their researchers. The problem is now the incentives are not increasing the number of subscriptions but increasing the number of papers published.
Only if the authors pay for it.
https://authors.acm.org/open-access/acm-open-for-authors-hom...
> From January 1, 2026, all ACM publications will be published Open Access (OA), free to read, share, and reuse in the ACM Digital Library.
So will LLMs have premium-tier access or basic-tier access?
it is open like open in openai
I'm really not sure how this would work. I don't know how the ACM works, but in IEEE you would have to give them your publishing rights. However, training a LLM is not publishing by itself, it is a derivative work? Any way, at this point authors should be entitled to monetary compensation, not the publisher. The deal is totally different.
They probably already scraped it.
Of course they did. There is open access to this library. These guys actually think they're offering new training data? It's kind of hilariously naive.
No doubt
Blocking access only hurts people who follow the rules. Unblocking access lets them compete with those who break the rules.
I think the right choice is pretty clear...
Would you prefer a parquet dump of acm articles to hugging face?
A llm emulating a person is why many of my used sites banned llms due to scraping bandwith costs
Is this about access or accessibility to claude (for example)
I am sure the entirely of human computing knowledge is not that big.
Well if LLMs get access to it, we humans should get free access to it as well!
ACM has been leaning heavily into AI-generated content for their journals in the past year, and this article is no exception: it appears to be 100% AI generated and full of LLM verbiage.
There's something hilarious about that, but also, snake eating its own tail.
I don't think this is AI-generated. It is focused and direct. It reads like anodyne albeit totally human academic manager writing.
Something something Roko's basilisk
Are they in the position to do that?
What about the authors?
The ACM sent around a nice query to members which made it clear that they were going to do it even if 100% of the members said "No, don't do that".
I suppose they could be sued.
Don't you grant ACM a right to distribute your work when publishing? So doesn't ACM already have the right to grant access to AI?
It's already in there
the digital library should have always been open access
now it will be fodder for the slop machine
(i think that LLMs are going to wreck the peer review system for all but hard-experimental papers)
Since starting to use AIs seriously for search in the last 3 months I have read and referenced more published papers and academic primary sources then I think I did in the previous 5 years. They're fantastic for pointing at some claim and asking for the primary source for it, and then it does the work of following things through the layers of backref to the original, assuming it's online. Opening the library up to AI access makes it far more accessible and usable then it was before.
AIs, at least in their current form, make you more who ever you were. If you want snap, glib answers of dubious accuracy, they'll give them to you, more easily than ever before. If you want to dig back into primary sources and get the original content, they'll do that for you, more easily than ever before.
Can't speak to how the science infrastructure is going to handle them, but if it takes down the peer review system, which I think has been worthless for probably going on two decades and has just given the entire enterprise a false sense of assurance, it'll probably be a net gain in the end. Peer review is a source of more problems then it is solving right now.
The real efficacy of the peer review system has always been somewhat questionable. That's a big reason why arxiv is so prominent.
Came to say the same thing. Now remains the time to give the public access to the ACM digital library.
ACM agrees I think: https://dl.acm.org/openaccess
> Beginning January 2026, all ACM publications and related artifacts in the ACM Digital Library will be made open access.
Wait, so did this already happen and it just didn’t get attention?
It was a front page story here on HN back in December 2025:
https://news.ycombinator.com/item?id=46313991
It got a lot of discussion here when it happened and before as they phased it in.
> I think that LLMs are going to wreck the peer review system for all but hard-experimental papers
I think I may be shadowbanned, but at what point do we start viewing LLMs as a national security threat?
ACM bureaucrats and AI shills are flagging dissent. Welcome to 2026. The entire Internet is just top down propaganda.
Now? That's cute.
TLDR: “we’re going to try to get money from LLM providers for access to our back catalog without getting permission from the authors or providing them with any share of the revenue”.
Scholarly article authors never expected royalty or permission for their scholarly works to be reused.
I don't want to be mean but if it's that hard to even recognize that LLMs are even a valid thing that could intersect with your business that you need some kind of campaign for it..
It's almost like, "don't hurt yourself unc, we will just search arxiv".