Unlock Unstructured Data in SageMaker Catalog with Generative AI Queries
This AWS Big Data blog post details the consumer-side workflow for accessing and querying unstructured data published to Amazon SageMaker Catalog, building on an earlier post about data ingestion and enrichment. It showcases how businesses can extract critical insights from diverse unstructured data types like PDFs and emails using generative AI. The solution leverages Amazon Bedrock for natural language queries, offering both a no-code chat agent for data analysts and programmatic access for application engineers. Governed access to these enriched assets is maintained through SageMaker Catalog's approval workflow.
- →Add a name: producerprojectdata
- →Add a name: MedicalKB
- →Add a name:
Features (6) ›
- Add a name: producerprojectdata
Add the producer’s S3 path as a new S3 location: s3://amzn-sagemaker-bucket-<domain-id>-<project-id>/medical/. Note: You can get the S3 location details from the technical name of your subscribed asset. - Choose the AWS Region, and then choose Add data to add this as a new location. Note: Make sure the AWS Region you select supports the Amazon Bedrock foundation models used later in this post. For a list of available models by Region, see Supported Regions and models for Amazon Bedrock. Note: Make sure to select only the PDF files within the S3 path for the data source
Add a name: After the location is added, it appears as a selectable S3 location when creating a knowledge base in AI Apps. Complete the following steps to configure the chat agent app:
- Add a name: MedicalKB
Add a description: Knowledge base built from subscribed medical S3 data assets. Contains medical documents used to provide grounded, context-aware responses to medical domain queries
- Add a name:
https://aws.amazon.com/blogs/big-data/query-unstructured-data-in-amazon-sagemaker-catalog-using-generative-ai/
Related releases
- SageMaker Unified Studio Supports IAM Permissions Boundaries for Tooling Blueprints AWS Big Data Blog ·
- Amazon SageMaker HyperPod Inference Gateway for Scalable LLM Inference AWS What's New ·
- Amazon EMR on EKS now supports Spark Connect for interactive workloads AWS What's New ·
- Amazon EMR on EKS Now Supports IPv6 for Scalable Big Data Workloads AWS What's New ·
- Amazon SageMaker Unified Studio Gains Domain-Level VPC Networking AWS Big Data Blog ·
- Spark Connect now supported on Amazon EMR on EC2 for interactive PySpark AWS Big Data Blog ·