Serverless AI and Data with Modal : Running GPU Workloads, Pipelines, and Apps Without Managing Infrastructure

"Serverless AI and Data with Modal: Running GPU Workloads, Pipelines, and Apps Without Managing Infrastructure"

For experienced Python engineers, ML practitioners, and platform-minded developers, this book offers a rigorous guide to building serious AI and data systems on Modal without operating clusters, schedulers, or container fleets. Rather than treating serverless as a convenience feature, it presents Modal as an execution model for production workloads—one that changes how you think about deployment boundaries, GPU access, autoscaling, and the relationship between code, runtime, and infrastructure.

Across the book, readers learn how to package reproducible runtime environments, deploy remote functions and persistent applications, tune concurrency and warm capacity, and run GPU-backed inference, training, and batch jobs with predictable performance. It also covers map-style distributed pipelines, scheduled automation, HTTP endpoints and framework-based APIs, as well as secrets, storage, observability, and failure-aware operations. The emphasis is on architectural judgment: when to choose denser container utilization over more replicas, how to avoid deployment-time packaging failures, and how to design systems that remain correct under retries, cold starts, and scaling pressure.

The book assumes strong Python fluency and comfort with modern software delivery and cloud concepts. Its distinguishing strength is depth: each topic is framed operationally, with close attention to version-sensitive guidance, trade-offs, and production realities that matter once prototypes b

Om den här boken

"Serverless AI and Data with Modal: Running GPU Workloads, Pipelines, and Apps Without Managing Infrastructure"

For experienced Python engineers, ML practitioners, and platform-minded developers, this book offers a rigorous guide to building serious AI and data systems on Modal without operating clusters, schedulers, or container fleets. Rather than treating serverless as a convenience feature, it presents Modal as an execution model for production workloads—one that changes how you think about deployment boundaries, GPU access, autoscaling, and the relationship between code, runtime, and infrastructure.

Across the book, readers learn how to package reproducible runtime environments, deploy remote functions and persistent applications, tune concurrency and warm capacity, and run GPU-backed inference, training, and batch jobs with predictable performance. It also covers map-style distributed pipelines, scheduled automation, HTTP endpoints and framework-based APIs, as well as secrets, storage, observability, and failure-aware operations. The emphasis is on architectural judgment: when to choose denser container utilization over more replicas, how to avoid deployment-time packaging failures, and how to design systems that remain correct under retries, cold starts, and scaling pressure.

The book assumes strong Python fluency and comfort with modern software delivery and cloud concepts. Its distinguishing strength is depth: each topic is framed operationally, with close attention to version-sensitive guidance, trade-offs, and production realities that matter once prototypes b

Kom igång med den här boken idag för 0 kr

  • Få full tillgång till alla böcker i appen under provperioden
  • Ingen bindningstid, avsluta när du vill
Prova gratis nu
Mer än 52 000 personer har gett Nextory 5 stjärnor i App Store och på Google Play.