Five years after Eliza Zhang’s book-scanning video reached millions of people, the Internet Archive’s page-by-page approach to digitization is capturing attention all over again.
In February 2021, a video of Internet Archive book scanner Eliza Zhang turning the pages of a book went unexpectedly viral. It reached 7.7 million people in its first 48 hours and became one of our most widely seen social posts ever.
Five years later, it happened again.
As renewed attention to the way books are digitized brought new questions about our own scanning practices, people once again discovered Eliza, the Scribe, and the decidedly hands-on process behind the digital books in our collections:
A big thank you to Eliza, and to all of the book scanners, digitization staff, librarians, engineers, archivists, and volunteers who spend their days caring for the physical record and making it accessible in digital form.
Rather than wait for a traditional publisher, the Berkeley, California, teen says she felt an urgency to get her message out and reach as many people as possible. On July 24, she released her second book, Porcelain Rage: A Manifesto for the Gen Z Woman, online at the Internet Archive.
In her new book, Elias reflects on cultural trends impacting young women and the challenges that come with the digital age. It explains the good, the bad, and the complexity of the moment, weaving in philosophy and history. The chapters include critiques of the Barbie movie, “sad girl” pop music, pilates, and podcasts.
She analyzes the pressures facing her generation and says Gen Z women are reclaiming femininity in a new way.
“This decade, we’ve decided to burn the playbook and make our own,” Elias writes in the forward to the book. “We can hold two truths simultaneously: the porcelain-perfect, pink image of the women of the past and the rageful, angry, sad, feral versions of ourselves. The contradiction is the point. It is our reality.”
Her new book also has a chapter about her struggle with mental health and chronicles the mistreatment of girls who have been institutionalized, in a parallel to Nellie Bly’s 19th-century expose Ten Days in a Madhouse. Through her writing, Elias says she wants to humanize a generation that she claims is often misunderstood.
“It’s about balancing the expectations of perfection in the modern age and the rage and mental illness that we are experiencing at catastrophic levels these days,” Elias says.
Register for a book talk with Elias and writer Sonia Mansfield on August 4.
Elias comes from a family of writers (her mother, Jessica Carew Kraft, author of Why We Need to Be Wild; her grandfather, Thomas Elias, Pulitzer-nominated columnist for California Focus; and her grandmother, Marilyn Elias, a journalist for USA Today). Her father, Jordan Elias, is a plaintiffs’ class-action lawyer who helped spearhead a case against Google for monopolizing digital advertising and advocates for the public’s right to access judicial records and histories to protect against sealing by corporations.
Simone grew up in San Francisco and attended preschool across the street from the Internet Archive at Congregation Beth Sholom. Since she was 13, she has been involved with the Internet Archive, conducting a series of podcast interviews, judging film contests, and helping with public-domain movie nights.
“When I was looking for a place to house a book that is so firmly rooted in public utility and using information for the improvement of people, it was obvious to go with [the Internet Archive],” Elias says.
Choosing the open-access route and bypassing a publisher, she says she was able to be authentic in her analysis. Her hope is that removing any barriers means a broad audience will read the book.
“The book was not so much an end to a conversation, but a start. The ultimate goal is to have young women talking about this with each other instead of holding it in,” Elias says. “It’s supposed to function as a manifesto, a survival guide, a definitive blueprint of what we’re dealing with in America as women today through our culture.”
In an introductory note about the publishing model for Porcelain Rage, Elias says her choice to post her book on archive.org may seem jarring or unprecedented. Yet she argues this avenue allows her freedom to tell her story unfiltered.
“I’m sick of how the industry puts women into neat little saleable categories,” Elias writes of traditional publishers in the book. “We don’t need more palatable stories in an unpalatable world. This book is free, uncensored, and completely yours. That’s how it should be.”
BOOK TALK Internet Archive will host a book talk with Elias and author Sonia Mansfield on August 4th at the Internet Archive. Registration is free. If you can’t make it in person, the discussion will be livestreamed. – Register for the in-person gathering at the Internet Archive. – Watch the livestream.
This year, as part of Google Summer of Code (GSoC), the Internet Archive is collaborating with two outstanding contributors to make it easier for patrons to find relevant books on Open Library.
Tanishq Sangwan, a 19-year-old from Gurugram, India, and Chisom Nnamani of Lagos, Nigeria, are two of 1,141 software developers from around the world who have been selected through GSoC to hone their engineering capabilities with open-source organizations.
“The Internet Archive’s focus for 2026 is: tools for participation. Participation must be earned by building an experience patrons want to return to.” said Mek, program lead for Open Library. “The work that both Chisom and Tanishq are doing is central to creating a more reliable Open Library experience, where readers can repeatedly discover, access, and enjoy books.”
Tanishq Sangwan
Sangwan, who just completed his second year of college studying artificial intelligence, is focusing on the journey of patrons who join Open Library and use it once, but do not return. His objective is to understand where patrons encounter barriers and identify opportunities to create a more useful, lasting experience.
Sangwan comes into the project with two years of experience from ZNotes, an educational organization that provides free notes and videos from students around the world.
“I built this passion and got this amazing feeling when my work was making an impact on people and they were receiving some value,” he said. “Now, Google Summer of Code is a wonderful opportunity to connect with open-source organizations and Open Library where its work directly impacts people’s lives.”
“This year’s collaboration is important because lots of patrons discover Open Library, but too often don’t always find the books they want,” Mek said.
The work begins when a patron lands on a book that is unavailable for reading. Soon, instead of reaching a dead end, patrons will be presented with nearby books on the same shelf that are available now.
When patrons find a relevant, available book, a simpler registration process will help them get started with fewer steps and return to the book they found. Furthermore, Sangwan is helping patrons connect with new book recommendations on an ongoing basis by introducing an activity feed to the account page.
“I like to hear about the patron psychology, how they’re interacting with the platform, and what’s going in their mind from the first moment to the very last,” he said. “We’ll be researching and conducting interviews with lost patrons so we can connect this bridge between patrons and the millions of books in our catalog.”
Chisom Nnamani
Nnamani, who already has certifications in Data Analytics and Data Engineering, just completed her sophomore year pursuing a second degree in Computer Science. At Open Library, Nnamani is leading a major cleanup effort to add structured tag data to books so they can be searched by genre and subgenre. Her work is paving the way for the addition of a wide variety of new searchable tags, including: moods, fiction and non-fiction, content warnings, and literary formats, such as memoirs, biographies, and more. By cleaning up messy data and enabling better genre and subject browsing, Chisom is helping remove barriers preventing patrons from discovering books they love.
“Today, many of our subject pages feel computer generated and can’t compare to the beautiful, curated experiences you find at small book stores,” says Mek. “The work Chisom is leading to map the messy subject tags we have to clear genres and subgenres will help us offer patrons a more useful and satisfying browsing experience.”
“I really care about books being accessible to people,” Nnamani said. “In Nigeria, we have limited access to physical libraries, so Open Library is something that matters. It gives everyone the opportunity to come and read any kind of book and gain insights.”
“This project is inspiring to me,” Nnamani said. “I like to work on projects where I can connect the data and infrastructure in ways that contribute to the organization’s goals. I enjoy helping to solve complex problems at the intersection of systems and data.”
Open Library Fellows work remotely, but meet regularly online with Internet Archive staff and mentors. At the end of the summer, each contributor will publish a blog post explaining their technical journey and experience gained.
Since 2005, Google’s Summer of Code has supported more than 23,000 students from 123 countries with stipends to receive mentorship and contribute 48 million lines of code to over 1,000 open source organizations worldwide.
Three of North America’s flagship library organizations have thrown their weight behind the movement to protect memory institutions’ digital rights.
The American Library Association (ALA), the Association of Research Libraries (ARL), and the Canadian Association of Research Libraries (CARL) just joined the Statement on Four Digital Rights for Memory Institutions Online. Together, they represent thousands of public and academic research libraries, as well as three of Canada’s federal and parliamentary libraries. Now, they stand with Our Future Memory’s global coalition of libraries, museums, and other cultural heritage organizations expressing the urgent need to protect memory institutions’ vital role in the digital age.
In endorsing the Statement, Katherine McColgan, manager of administration and programs for CARL, explained that “[t]he current digital landscape is significantly affecting the knowledge economy in two ways. One is that online materials are on platforms that restrict the collection, preservation, and making available materials for future generations. The second is that, without the ability to digitize and make available important scholarly works online, information is lost to new generations of scholars. It is imperative that memory institutions are able to continue their work in the digital environment in the same way as with print.”
Indeed, the Statement demands nothing new—only the basic rights necessary for libraries, archives, museums and other cultural heritage organizations to continue their core operations and fulfill their public-serving mission. The Statement calls on policymakers around to world to ensure that memory institutions have the right and ability to:
Collect digital materials
Preserve digital collections
Provide controlled digital access
Cooperate across institutions
Building on well over a decade of advocacy by leaders in the library community, “[t]he statement’s principles provide policymakers with a clear roadmap for how to maintain the essential public role of libraries, archives, and museums in the digital age,” said Lisa Varga, associate executive director of ALA’s Public Policy and Advocacy Office.
It “underscores the importance of protecting libraries’ rights through legislative advocacy and licensing strategies, in an era of increasingly restrictive licensing agreements that threaten essential library functions like building collections, preserving materials, and enabling advanced computational research methods such as AI,” explained ARL’s director of public policy, Katherine Klosek.
With these new signatories, the global call to protect the rights of memory institutions online gains even further momentum.
Ready to Join?
Your organization can join the movement and sign the Statement by going to the Our Future Memory website.
The movement for Our Future Memory is getting bigger, with yet another library leader endorsing memory institutions’ digital rights.
The Boston Library Consortium (BLC), comprised of twenty-six research libraries in the New England area, has signed the Statement on Four Digital Rights for Memory Institutions Online. BLC joins more than forty other signatories from around the world, from the Wikimedia Foundation to the International Federation of Library Associations and Institutions (IFLA). Alongside two other recent signatories, it adds a strong voice to the growing list of libraries, museums, and other cultural heritage organizations that are calling for stronger legal protections to fulfill their public missions in the digital age.
“BLC is proud to join institutions worldwide in defending our rights to collect, preserve, provide access, and cooperate. Libraries safeguard cultural memory—and online content shouldn’t be an exception.”
Charlie Barlow, executive director, Boston Library Consortium
“BLC is proud to join institutions worldwide in defending our rights to collect, preserve, provide access, and cooperate,” said executive director Charlie Barlow. “Libraries safeguard cultural memory—and online content shouldn’t be an exception.”
In putting its name to the statement, BLC offers further proof that libraries and archives know exactly what they need to keep preserving and providing access to the culture record.
The global campaign to secure digital rights for libraries and memory institutions just gained a powerful new ally.
As explained in a post by Beatrice Murch of Internet Archive Europe, the International Federation of Library Associations and Institutions (IFLA)—the leading international body representing the interests of library and information services—has signed the Statement on Four Digital Rights of Memory Institutions, joining more than 30 signatories from around the world who are calling for the legal rights that libraries, archives, and other cultural heritage organizations need to fulfill their missions in the digital age.
It’s such a good initiative. I think as far as we were concerned, when we looked [at] the Four Digital Rights […], we sat down and thought this stuff is obvious, isn’t it? This is just reaffirming the things that libraries have always done.
These are basic functions that need to be in place, not just to deliver library rights, but ultimately library rights are the rights of the community that depends on libraries to actually get things done, to fulfill their own rights, to fulfill their own potential.
Stephen Wyber, IFLA
In joining the statement, IFLA strengthens the growing international movement to secure the legal foundations for long-term digital preservation and access to knowledge. Their endorsement signals that libraries and archives worldwide are aligned in calling for legal reform on four essential rights:
A recent legal decision has reaffirmed the power of fair use in the digital age, and it’s a big win for libraries and the future of public access to knowledge.
On June 24, 2025, Judge William Alsup of the United States District Court for the Northern District of California ruled in favor of Anthropic, finding that the company’s use of purchased copyrighted books to train its AI model qualified as fair use. While the case centered on emerging AI technologies, the implications of the ruling reach much further—especially for institutions like libraries that depend on fair use to preserve and provide access to information.
What the Decision Says
In the case, publishers claimed that Anthropic infringed copyright by including copyrighted books in its AI training dataset. Some of those books were acquired in physical form and then digitized by Anthropic to make them usable for machine learning.
The court sided with Anthropic on this point, holding that the company’s “format-change from print library copies to digital library copies was transformative under fair use factor one” and therefore constituted fair use. It also ruled that using those digitized copies to train an AI model was a transformative use, again qualifying as fair use under U.S. law.
This part of the ruling strongly echoes previous landmark decisions, especially Authors Guild v. Google, which upheld the legality of digitizing books for search and analysis. The court explicitly cited the Google Books case as supporting precedent.
While we believe the ruling is headed in the right direction—recognizing both format shifting and transformative use—the court factored in destruction of the original physical books as part of the digitization process, a limitation we believe could be harmful if broadly applied to libraries and archives.
What It Means for Libraries
Libraries rely on fair use every day. Whether it’s digitizing books, archiving websites, or preserving at-risk digital content, fair use enables libraries to fulfill our public service missions in the digital age: making knowledge available, searchable, and accessible for current and future generations.
This decision reinforces the idea that copying for non-commercial, transformative purposes—like making a book searchable, training an AI, or preserving web pages—can be lawful under fair use. That legal protection is essential to modern librarianship.
In fact, the court’s analysis strengthens the legal groundwork that libraries have relied on for years. As with the Google Books decision, it affirms that digitization for research, discovery, and technological advancement can align with copyright law, not violate it.
Looking Ahead
This ruling is an important step forward for libraries. It reaffirms that fair use continues to adapt alongside new technologies, and that the law can recognize public interest in access, preservation, and innovation.
As we navigate a rapidly changing technological landscape, it’s more important than ever to defend fair use and support the institutions that bring knowledge to the public. Libraries are essential infrastructure for an informed society, and legal precedents like this help ensure they can continue their vital work in the digital age.
Louis Brizuela says managing the microfiche digitization center for Democracy’s Library gives him a sense of pride. “I feel like I’m making a difference,” said the 28-year-old who lives in the Bay Area. “We’re scanning and preserving all this really cool content.”
Brizuela and his six-person team are currently digitizing U.S. Supreme Court case documents and government records from Canada dating back to the 1930s. The documents are stored on microfiche cards, a flat, film-based format commonly used from the mid-20th century for preserving and accessing paper records, which requires a specialized reader for viewing—making the information contained on the cards difficult to access. “It’s useful for law students or anybody – and it’s free to use without borders,” he said. “Also, it’s valuable for the sake of archiving so information doesn’t get lost.” Next, Brizuela said he’s looking forward to receiving a donated collection of microfiche with images of Sanskrit Buddhist tablets.
Anyone can watch the crew in action on a livestream of the microfiche scanning operation (https://www.youtube.com/watch?v=wG9hiWSrSD4). Activity occurs Monday–Friday, 7:30am-3:30pm and 4:00pm-midnight U.S. Pacific Time (GMT+8)—except U.S. holidays. Mellow lo-fi music plays in the background during working hours and continues with various video and still images from the Internet Archive’s collections rotating on the feed when the digitization center is closed.
During the livestream, one camera is focused on an operator feeding microfiche cards beneath a high-resolution camera; another other provides a close-up view of the material. Each page is processed, made fully text-searchable, and added to the Internet Archive’s public collections. Researchers and readers can easily access and download the documents freely through Democracy’s Library.
Brizuela said the staff has embraced the public window on their work. He joined the Internet Archive in February and hired people who were willing to be on camera and understood the potential benefit of the exposure. “It’s not like ‘Oh, Big Brother is watching’,” he said, noting the employees have fun with the situation. “We’re not robots. We do show our characters. We’re human.”
The team is leaning in, Brizuela said, suggesting they dress up in costumes for Halloween and maybe wearing elf hats at Christmas to add a festive touch to the project. They also answer questions in a live chat with viewers.
Brizuela comes to this position from a varied career working in the military, medical fields, retail and web development. He’s long had an interest in photography, particularly shooting and developing his own 35mm film. So, Brizuela said, it was not hard to pick up how to operate the custom-built scanner and oversee the digitization process.
Louis Brizuela stands in front of a custom-built microfiche scanning workstation.
Every morning, the team huddles up in the small digitization center to talk about the previous day’s completed pages and map out the upcoming work. Brizuela watches over and QA’s the scanning done by the team. Depending on the type of collection, each scanner can scanhundreds of cards a day.
Brizuela describes the vibe in the microfiche digitization center as pretty relaxing, with staff members chatting and interacting while they work. Often, they have headphones to listen to an audiobook or podcast. “If they are listening to music, sometimes they bust a dance move, or bob their head to get in the groove. People enjoy seeing that,” Brizuela said.
How is knowledge created, shared, and preserved in the digital age—and what forces are shaping its future?
We’re thrilled to announce the launch of Future Knowledge, a new podcast from the Internet Archive and Authors Alliance. Hosted by Chris Freeland, librarian at the Internet Archive, and Dave Hansen, executive director of Authors Alliance, the series brings together authors, librarians, policymakers, technologists, and artists to explore how knowledge, creativity, and policy intersect in today’s fast-changing world.
In each episode, an author discusses their book or publication and the big ideas behind it—paired with a thought-provoking conversation partner who brings a fresh perspective from the realms of policy, technology, libraries, or the arts.
We’re kicking off the podcast with a double feature—two episodes tackling copyright history and AI’s global impact:
Episode 1: The Copyright Wars
Historian Peter Baldwin joins copyright scholar Pamela Samuelson to unpack The Copyright Wars—a sweeping look at 300 years of trans-Atlantic copyright battles. From 18th-century publishing monopolies to today’s clashes between Big Tech, libraries, and the entertainment industry, this conversation reveals how history can illuminate the future of intellectual property in a digital world.
Episode 2: Copyright, AI, and Great Power Competition
Authors Joshua Levine and Tim Hwang sit down with Lila Bailey to discuss Copyright, AI, and Great Power Competition. Together they explore how artificial intelligence is transforming copyright law—and how global powers are using IP policy as a strategic tool in the race for technological dominance.
Whether you’re an author thinking about how to share your work, a librarian navigating digital access, or a curious listener exploring how knowledge shapes our world, Future Knowledge is for you.
For many years, the Drug Policy Alliance (DPA) maintained a large library of books on drug use and policy at its New York City headquarters. As researchers shifted to working online, DPA’s Jules Netherland said she noticed fewer people coming into the office to use the collection.
“It became clear if we really wanted people to benefit from our resources that digitization was the way to go,” said Netherland, managing director of the Alliance’s Department of Research and Academic Engagement. It was also an opportunity to add to the growing collection of the Substance Abuse Librarians and Information Specialists (SALIS).
DPA donated its book collection to the Internet Archive to be digitized and made available for lending and for the print disabled. A team was sent to New York to pick up the books, which were packaged onto three pallets and shipped to a facility for scanning and storage.
Now, the digital version of the DPA library, with 2,260 items, is available to the public at https://archive.org/details/dpa. It is part of the larger SALIS collection of 8,647 items on alcohol and substance abuse digitized by SALIS.
The public has responded with curiosity. In January, 10,000 items were accessed in the digitized collection. Rosenberg speculates the audience is likely researchers, historians, healthcare providers, and policymakers.
In the rapidly evolving field of drug policy, which spans many disciplines, Netherland said it’s important to provide evidence-based information to the public. The hope is to enhance advocacy efforts with easier access to the organization’s collection. DPA developed a resource guide to encourage its use on the Internet Archive.
In donating its collection, DPA helped build the Internet Archive’s SALIS collection. Since 2008, SALIS has helped preserve thousands of items from physical libraries with research from drug and alcohol fields that have closed, said Andrea Mitchell, SALIS executive director.
About 30 years ago, there were approximately 95 libraries, clearinghouses, and resource centers around the world devoted to collecting, cataloguing, and disseminating information concerning alcohol, tobacco, and other drugs, Mitchell said. However, today the majority of those libraries or databases have closed. The U.S. government has also shut down collections, including the National Institute on Drug Abuse, whose library went back to 1935. “We’re losing important resources and knowledge,” Mitchell said.
This leaves a void in access that has been filled, in part, by digitized collections online. Mitchell said The SALIS Collection includes materials that go back to 1774 and books from medicine, sociology, psychology, economics, law and policy, criminal justice, and other fields. In addition to books, there are government documents, grey literature, and newsletters.
The DPA collection was one of the larger libraries in the U.S., Mitchell said, and its donation to the Internet Archive is significant and welcome.
The Internet Archive is interested in receiving more curated collections like DPA’s on specific subject matters, Rosenberg added. “These really valuable books for research and resources are often not preserved when funding is lost at the library that houses them,” she saidTo find out more about the physical item donation process, go to the Help page for details.