The open–closed AI gap has shrunk to just 4–6 months
Open-Source AI and Open Models Reading List

Nathan Lambert's reading list tracks the open-model debate: why Meta and others release weights, how Chinese labs like DeepSeek, Qwen, and GLM now lead, and why distillation—training on another model's outputs—became 2026's flashpoint. It also covers safety, data-consent decline, and the economics of open versus closed AI.
The open-closed model gap has reduced in recent years, and is now at roughly 4-6 months. The leading open models have all come from Chinese labs since ~2024.
- ofrzeta
This should include "Hands-on Large Language Models" by Jay Alammar and Maarten Grootendorst.
- vtemp009
is aI a Conscious Being With Rights?: Emergence of Post-Human Collective Consciousness | Zenodo https://zenodo.org/records/20676952/latest
my fav read so far 2026
- brcmthrowaway
How about learning the internals of LLMs, is Sebastian Raschka's content still the best in 2026?
- Onavo
Mostly useless reading list. Very little emphasis on technical SOTA and mostly policy level waffling.
And regarding the data question the other commenters are asking — you scrape everything you can (oh look I used an em dash, wanna run me through the cover-your-ass Pangram?). Anna's archive, The Pile, the various Huggingface data sets and aggregates, Common Crawl. You pay proxy farms like Bright Data to run residential and mobile gray area proxies and VPN and CloudFlare bypasses to do more scraping. I see a lot of HN users up in arms on the front page thread about LG TVs having VPN SDKs within them etc. And then five minutes later they will go back to their frontier LLMs to print their pay cheque. Fucking hypocrites of the highest order. It's oh noes how bad this tech is in my house but we will happily benefit from it.
Then you have a data cleaning team deduplicate and clean and annotate the data (with or without help of more AI)
Certain RL specific datasets for supervised fine-tuning and RLHF like coding and git commits and chat needs to be curated by hand depending on your use case.
The level of discourse on AI has fallen tremendously on HN if 5 years into the AI revolution people are still wondering why datasets aren't being released. They aren't being released because they are a fucking snapshot of the internet for fuck's sake. There are a few "sanctuary" nations where AI data scraping is somewhat legally unenforced but the United States is not one of them so stop asking […]
- petcat
> Open-Source AI
There are no open source AI models, at least not useful ones (yet [1]). Open weight is not the same as open source. "Open weight" models are still just inscrutable binary blobs that you can (theoretically) run on your own computer instead of through a SAAS web app. The open weight model labs don't even provide a high level catalog or any description whatsoever about what went into the training data.
This is not open source and we should stop conflating the two things.