Scribd leverages Gemini Enterprise batch inference for 400M+ document classification
Scribd successfully utilized Gemini Enterprise with batch inference to classify over 400 million user-generated documents, spanning 12 billion pages, for trust and safety purposes. This implementation allowed a corpus-wide backfill in months, processing more than 99% of PDFs natively without complex pre-processing pipelines. Gemini's multimodal understanding and batch pricing, offered at a 50% discount to interactive rates, made large-scale LLM classification economically viable. The project, a collaboration with Google Cloud, now forms the foundation for continuous content evaluation of newly uploaded materials.
- →Large-scale content classification challenge at Scribd
- →Gemini's native PDF understanding and multimodal capabilities
- →Cost-effective batch prediction for corpus-wide backfill
- →Collaborative partnership with Google Cloud and future plans
Notes (4) ›
- Large-scale content classification challenge at Scribd
Scribd needed to perform trust and safety classification across its 400M+ document corpus, comprising over 12 billion pages, but found traditional methods and existing moderation tools inadequate. They sought a single, cost-effective solution capable of multimodal understanding across diverse content types, languages, and contexts.
- Gemini's native PDF understanding and multimodal capabilities
Gemini's ability to accept PDFs natively and process each page as both text and image was a key differentiator, eliminating the need for OCR or rendering pipelines for over 99% of Scribd's content. This multimodal approach allowed Gemini 2.5 Flash Lite and Pro models to identify visual policy signals that text-only moderation endpoints often missed.
- Cost-effective batch prediction for corpus-wide backfill
Gemini Enterprise's batch prediction model, priced at a 50% discount compared to interactive rates, made large-scale LLM classification economically feasible. The execution model was simplified: documents were staged in Cloud Storage, submitted to Gemini, and results flowed back into Scribd's data platform, avoiding serving infrastructure overhead.
- Collaborative partnership with Google Cloud and future plans
Google Cloud partnered with Scribd to plan the backfill, advise on region strategy, and scale batch throughput, resulting in faster-than-projected job completion. This successful project now underpins continuous evaluation of newly uploaded content and is being extended to other content-understanding workloads, significantly streamlining future roadmap items.
https://cloud.google.com/blog/topics/customers/scribd-inc-classifies-millions-of-documents-on-gemini-enterprise/
Related releases
- GCP Storage Intelligence Advisor Reaches GA, Expands Batch Operations Google Cloud Blog ·
- Google's GKE & Cloud Run gain AI optimizations, new security features; cited as Gartner Leader Google Cloud Blog ·
- GKE agentic migration open-sourced for AI-assisted EKS-to-GKE Kubernetes migrations Google Cloud Blog ·
- Google open-sources AI-assisted plugin for EKS-to-GKE migrations Google Cloud Blog ·
- Latin American Midsize Businesses Drive Digital Transformation with Google Cloud AI Google Cloud Blog ·
- Google Cloud details why AI startups choose its comprehensive stack Google Cloud Blog ·