gcp Google Cloud Blog ·

Scribd leverages Gemini Enterprise batch inference for 400M+ document classification

blogaigcpgaengineergcp-cloud-storage
announcement

Scribd successfully utilized Gemini Enterprise with batch inference to classify over 400 million user-generated documents, spanning 12 billion pages, for trust and safety purposes. This implementation allowed a corpus-wide backfill in months, processing more than 99% of PDFs natively without complex pre-processing pipelines. Gemini's multimodal understanding and batch pricing, offered at a 50% discount to interactive rates, made large-scale LLM classification economically viable. The project, a collaboration with Google Cloud, now forms the foundation for continuous content evaluation of newly uploaded materials.

  • →Large-scale content classification challenge at Scribd
  • →Gemini's native PDF understanding and multimodal capabilities
  • →Cost-effective batch prediction for corpus-wide backfill
  • →Collaborative partnership with Google Cloud and future plans
Notes (4) ›
  • Large-scale content classification challenge at Scribd

    Scribd needed to perform trust and safety classification across its 400M+ document corpus, comprising over 12 billion pages, but found traditional methods and existing moderation tools inadequate. They sought a single, cost-effective solution capable of multimodal understanding across diverse content types, languages, and contexts.

  • Gemini's native PDF understanding and multimodal capabilities

    Gemini's ability to accept PDFs natively and process each page as both text and image was a key differentiator, eliminating the need for OCR or rendering pipelines for over 99% of Scribd's content. This multimodal approach allowed Gemini 2.5 Flash Lite and Pro models to identify visual policy signals that text-only moderation endpoints often missed.

  • Cost-effective batch prediction for corpus-wide backfill

    Gemini Enterprise's batch prediction model, priced at a 50% discount compared to interactive rates, made large-scale LLM classification economically feasible. The execution model was simplified: documents were staged in Cloud Storage, submitted to Gemini, and results flowed back into Scribd's data platform, avoiding serving infrastructure overhead.

  • Collaborative partnership with Google Cloud and future plans

    Google Cloud partnered with Scribd to plan the backfill, advise on region strategy, and scale batch throughput, resulting in faster-than-projected job completion. This successful project now underpins continuous evaluation of newly uploaded content and is being extended to other content-understanding workloads, significantly streamlining future roadmap items.

Read the original announcement →

https://cloud.google.com/blog/topics/customers/scribd-inc-classifies-millions-of-documents-on-gemini-enterprise/

Related releases