Articles liés à Multimodal Data Engineering: Build Production-Ready...

Multimodal Data Engineering: Build Production-Ready AI Data Systems with Python, Data Pipe-lines, Vector Databases, Multimodal RAG, Streaming, and Cloud Architecture - Couverture souple

Keldran, Soren

 
9798173578624: Multimodal Data Engineering: Build Production-Ready AI Data Systems with Python, Data Pipe-lines, Vector Databases, Multimodal RAG, Streaming, and Cloud Architecture

Synopsis

AI applications are moving beyond text—but most data pipelines were never designed for what comes next.

Getting a multimodal AI prototype to work is one thing. Engineering the infrastructure that keeps documents, images, audio, video, structured data, metadata, embeddings, and retrieved context synchronized and trustworthy in production is another.

You may already know Python, SQL, ETL, APIs, databases, or cloud fundamentals. But modern AI data engineering introduces new challenges. How do you build reliable AI data pipelines without losing provenance? Keep embeddings and vector indexes synchronized when sources change? Combine batch processing with real time data pipelines? Prevent stale or unauthorized evidence from reaching production RAG systems?

Multimodal Data Engineering gives you the architecture, engineering principles, and practical techniques to solve these problems in production.

Through one evolving Production Multimodal Data Platform, you'll learn how modern AI data architecture connects ingestion, storage, processing, metadata, representations, retrieval, streaming, governance, and operations.

Inside, you'll learn how to:

  • Engineer pipelines for text, PDFs, images, audio, video, and structured data

  • Design reliable batch, streaming, and event-driven architectures

  • Build object storage, lakehouse, metadata, lineage, and versioning layers

  • Process multimodal data with OCR, transcription, segmentation, enrichment, and quality controls

  • Engineer version-aware embedding pipelines and apply vector database engineering principles

  • Build semantic search, hybrid retrieval, metadata filtering, and reranking

  • Develop Kafka-based streaming with ordering, replay, backpressure, idempotency, and recovery

  • Engineer multimodal RAG with evidence retrieval, context assembly, provenance, freshness, and evaluation

  • Apply security, governance, privacy, lineage, and AI data observability

  • Practice AI infrastructure engineering with Docker, Kubernetes, AWS, Azure, and Google Cloud

  • Optimize CPU, GPU, storage, network, retrieval, and infrastructure costs

  • Troubleshoot stale embeddings, index drift, consumer lag, retrieval failures, bad context, OOM failures, and compound system failures

  • Engineer reliability, disaster recovery, CI/CD, SLOs, incident response, and capacity planning


Designed with a Beginner → Professional progression, the book is accessible to ambitious readers with basic Python or general software and data knowledge while developing the architectural reasoning required for professional systems.

Whether you're a Data Engineer, AI Engineer, ML Engineer, MLOps Engineer, Software Engineer, Cloud or Platform Engineer, or Data Scientist, you'll learn how the pieces fit together—not just how individual tools work.

This is not another book about prompting or training foundation models.

It's about engineering the data systems that make multimodal AI reliable, searchable, scalable, governable, and production-ready.

If you're ready to move beyond traditional ETL and fragile AI prototypes, get Multimodal Data Engineering and start building production-ready AI data systems.

Les informations fournies dans la section « Synopsis » peuvent faire référence à une autre édition de ce titre.