The Church That Holds the Internet
Just blocks from the Presidio of San Francisco, where the Golden Gate Bridge meets the city, there is an old church that most people walk past without knowing what it holds. Inside those walls renovated quietly over the years to house servers more than pews nearly all of the internet's history is kept. Every website that has ever disappeared. Every page that was edited and then changed beyond recognition. Every digital artifact that would otherwise be gone forever, erased by time, neglect, or corporate decisions made in distant boardrooms.
This is the home of the Internet Archive, the San Francisco nonprofit that in October 2025 marked a once-in-a-generation achievement: its Wayback Machine had now preserved more than one trillion web pages. The city of San Francisco responded by officially declaring October 22 to be "Internet Archive Day." It was a rare moment when a quiet infrastructure project one most people never think about until they desperately need it received something like a ticker-tape parade.
The story behind that milestone is the story of one man's long obsession with memory, permanence, and the radical idea that the internet could become the greatest library ever built if someone was willing to do the exhausting, unglamorous work of actually building it.
A Singular Focus Since 1996
Brewster Kahle did not set out to become the archivist of the internet. His background was in artificial intelligence research, and in the early days of the web, he was searching for ways to organize information at scale. What he found instead was a problem no one else seemed willing to name: the internet was not preserving itself. Websites went dark. News organizations shut down or were acquired and their archives vanished. Companies deleted their own content with little notice. The web was growing faster than anyone could track, and there was no plan for what would happen to all of it.
In 1996, Kahle founded the Internet Archive as a passion project. The ambition was simple and enormous at the same time: universal access to all knowledge. In the decades since, the organization has grown into a repository that preserves 99+ petabytes of data the books, web pages, music, television, government information, and software of our cultural heritage and works with more than 1,200 library and university partners to create a digital library accessible to all.
The organization is known for the Wayback Machine, which lets users search the history of almost one trillion web pages. But its collection extends far beyond websites. It archives images, software, video and audio recordings, documents, and contains dozens of resources and projects that fill a variety of gaps in cultural, political, and historical knowledge. The breadth of what is stored inside the Internet Archive is almost impossible to comprehend and that is precisely the point.
As the Electronic Frontier Foundation noted in a September 2025 podcast episode examining the organization, "Access to knowledge not only creates an informed populace that democracy requires, but also gives people the tools they need to thrive. And the internet has radically expanded access to knowledge in ways that earlier generations could only have dreamed of so long as that knowledge is allowed to flow freely."
The Wayback Machine: Snapshots of a Changing Web
The Wayback Machine, launched in 2001, remains the organization's most widely known tool. It operates by taking snapshots of websites at regular intervals, building a chronological record of how the web looked and changed over time. Anyone can visit archive.org and enter a URL to see what a website looked like five, ten, or twenty years ago. Politicians cannot quietly delete old statements. Companies cannot pretend their products never had certain features. Journalists and researchers can trace the evolution of public information in real time.
Mark Graham, director of the Wayback Machine, has spoken at length about the forces that make this work necessary. "There are numerous incentives to put content online," he told the BBC in 2024. "But there's little pushing companies to maintain it over the long term." The risks, he explained, are manifold: technology fails, institutions fail, companies go out of business, news organisations are gobbled up by other organizations or shut down entirely. The Wayback Machine exists because none of those risks are hypothetical. They happen constantly.
The risks are manifold. Not just that technology may fail, but that certainly happens. But more important, that institutions fail, or companies go out of business. Mark Graham, director of the Wayback Machine, speaking to the BBC
Research published in 2024 underscored the scale of the problem. Studies showed that 25% of web pages posted between 2013 and 2023 have already vanished. For anyone trying to understand how the world looked during that decade researchers, journalists, historians, lawyers, regulators a quarter of the record is simply gone. The Internet Archive and a handful of similar organizations are, as the BBC described it, "the only things standing in the way of digital oblivion."
What One Trillion Pages Actually Means
The number sounds abstract until you try to do something without it. Imagine looking for a news story that ran on a local newspaper's website in 2017 only to find the publication folded in 2021 and the site is gone. Imagine searching for an academic paper hosted on a university server that was decommissioned during a budget cuts. Imagine trying to document what a government agency's website said about a policy during a particular administration, only to discover the page was edited and the old version never saved anywhere.
These are not edge cases. They are daily occurrences. The Wayback Machine handles more than 800,000 requests per day from users who need exactly this kind of access. Library partners more than 1,200 of them rely on the Archive's collections for their own research and preservation work. The trillion-page milestone represents not just an engineering achievement but a vast, public record of our digital lives, safeguarded for the future.
As the Library Journal's infoDOCKET noted when reporting the milestone in October 2025, the celebration drew internet pioneers Vint Cerf and Sir Tim Berners-Lee the inventor of the World Wide Web himself for video appearances. Cultural voices like Annie Rauwerda of Depth of Wikipedia and Luca Messarra of Vanishing Culture participated. The Internet Archive presented its 2025 Hero Award to Berners-Lee, recognizing his role in inventing the web and his enduring advocacy for an open, accessible internet. The city of San Francisco held a public rally on the steps of City Hall. A behind-the-scenes tour of the physical archive in Richmond, California gave guests a rare look at preservation labs and the journey of physical materials from donation to long-term public access.
It was, in every sense, a celebration of infrastructure the kind that usually goes unnoticed until it is gone.
The Cost of Preservation: Legal Battles and Lost Books
The October 2025 celebration came only after years of bruising legal fights that nearly ended the project. The Internet Archive had scanned and lent digital copies of books under a principle called Controlled Digital Lending essentially treating its digital collections the same way physical library books are lent, with the same number of copies available as the organization owned. Publishers sued, arguing this model violated copyright. The organization fought the case for years.
In November 2025, Ars Technica reported on the outcome. "We survived," Brewster Kahle told the publication. "But it wiped out the Library." The organization was forced to remove more than 500,000 books from its Open Library. The legal fights are over, for now, and the Archive faces no major lawsuits and no active threats to its collections. But Kahle said the losses changed what the institution could offer. "The world became stupider," he said, "when the Open Library was gutted."
The experience left scars and lessons. The organization has rebuilt its approach to digital lending, working more closely with publishers and rights holders. It has also diversified its legal and operational strategies, determined not to find itself in the same vulnerable position again. Kyle Courtney, a copyright lawyer and librarian who leads the nonprofit eBook Study Group, which helps states update laws to protect libraries, has described Kahle's vision in terms that are both modest and audacious: transforming the Internet Archive into a digital Library of Alexandria "but with a better fire protection plan," he joked.
The Physical Archive and the Human Work Behind the Data
For all the talk of petabytes and trillion pages, the Internet Archive is not purely a technology project. It has a physical presence rows of servers, yes, but also a building in Richmond, California where books arrive every day by the truckload, are catalogued, digitized, and preserved. Volunteers and staff handle materials that might otherwise end up in landfills. The organization accepts donations from individuals, institutions, and publishers who want their collections to have a permanent home.
During the October 2025 celebrations, the Archive opened its physical operations to the public for "Doirs Open 2025" a rare chance to walk through preservation labs, see rare acquisitions, and understand the journey of physical materials from donation to long-term public access. For many visitors, it was the first time they understood that behind every Wayback Machine search result, there is a team of people making decisions about what to save, how to save it, and how to make it accessible without violating anyone's rights.
The archive also holds government documents, microfilm, and materials from partners who lack the infrastructure to preserve them independently. As Sen. Alex Padilla (D-Calif.) noted when designating the Internet Archive a federal depository library in 2025, the organization is a "perfect fit" to expand access to federal government publications in an increasingly digital landscape. The designation gives the Archive formal recognition as an institution that serves the public interest in access to official information a role libraries have played for centuries, translated into the digital age.
What This Means for Lnk2It Readers
For anyone who curates links, builds resource collections, or helps others find information online, the Internet Archive is not just an interesting institution it is a foundational tool. When a link goes dead, the Wayback Machine is often the only way to recover what was there. When a source disappears from the open web, the Archive's snapshots may be the only surviving evidence. When researching the history of an idea, a product, a policy, or a publication, the Wayback Machine provides the temporal dimension that the live web cannot.
Understanding how the Internet Archive works its scale, its limitations, its legal history, its partnerships helps anyone who relies on web-based research make better decisions about what to trust, what to save locally, and how to plan for the inevitable moments when the live web fails to deliver what you need.
The Archive's one-trillion-page milestone is also a reminder that preservation is an ongoing act, not a finished state. The web keeps growing. Pages keep disappearing. The organizations doing this work operate on limited resources, face constant legal uncertainty, and depend on public support and institutional partnerships to continue. The tools exist because people chose to build them and they will only continue to exist if that choice is renewed, over and over.
Where the Archive Goes From Here
Despite the losses of the legal battles, Kahle and his team are not slowing down. The organization has emerged from its court fights with a clearer sense of what it can and cannot do and with new ideas about how to expand access within those boundaries. Partnerships with libraries and universities continue to grow. New digitization projects are underway. The organization is exploring ways to preserve more video, audio, and ephemeral content that typically vanishes faster than text-based web pages.
At the October 2025 celebration, Kahle spoke about the work ahead in terms that were characteristically modest and characteristically ambitious. The Wayback Machine hitting one trillion pages was not an ending. It was a marker on a much longer road proof that the work could be done at scale, and proof that it needed to continue. The web being built today will be the historical record of tomorrow. Whether that record survives depends on decisions being made right now, by people willing to do the work.
As Lawrence Lessig told the New York Times back in the early days of the Wayback Machine, Kahle was "defining the public domain" online. That work continues. It is messier, more contested, and more necessary than ever.
Why This Matters
Every link you have ever followed that led somewhere unexpected an old version of a website, a document that has since been updated, a source that no longer exists on the live web exists because someone, somewhere, decided to save it. The Internet Archive is the largest-scale expression of that decision in the history of the web. Its founder spent thirty years building something that most people take for granted and few people fully understand.
The trillion-page milestone is an achievement worth understanding not because of the number itself, but because of what it represents: a running argument that memory matters, that access matters, that the record of how we lived online is worth preserving even when it is inconvenient, even when it is contested, even when no one is paying attention.
The old church in San Francisco holds that argument in its servers. It is still being made, every day, one page at a time.
Where to Read Further
- Inside the old San Francisco church that houses nearly all of the internet's history CNN Business's November 2025 feature on the Archive's facilities and the trillion-page milestone, including video of the Presidio-adjacent building
- Building and Preserving the Library of Everything the Electronic Frontier Foundation's September 2025 podcast episode featuring Brewster Kahle in conversation with EFF's Cindy Cohn and Jason Kelley
- We're losing our digital history. Can the Internet Archive save it? the BBC's September 2024 deep dive into digital entropy, Mark Graham's perspective on the Wayback Machine, and the stakes of losing 25% of the web's 2013-2023 record
- Internet Archive's legal fights are over, but its founder mourns what was lost Ars Technica's November 2025 reporting on the outcome of the copyright battles, Brewster Kahle's reflections on losing 500,000 books from the Open Library, and the organization's path forward
- Milestones: Internet Archive Celebrates One Trillion Web Pages Preserved the Library Journal infoDOCKET's coverage of the October 2025 celebration, including details on Internet Archive Day, the Hero Award to Sir Tim Berners-Lee, and the city's public rally



