Serverless AI and Data with Modal : Running GPU Workloads, Pipelines, and Apps Without Managing Infrastructure

"Serverless AI and Data with Modal: Running GPU Workloads, Pipelines, and Apps Without Managing Infrastructure"

For experienced Python engineers, ML practitioners, and platform-minded developers, this book offers a rigorous guide to building serious AI and data systems on Modal without operating clusters, schedulers, or container fleets. Rather than treating serverless as a convenience feature, it presents Modal as an execution model for production workloads—one that changes how you think about deployment boundaries, GPU access, autoscaling, and the relationship between code, runtime, and infrastructure.

Across the book, readers learn how to package reproducible runtime environments, deploy remote functions and persistent applications, tune concurrency and warm capacity, and run GPU-backed inference, training, and batch jobs with predictable performance. It also covers map-style distributed pipelines, scheduled automation, HTTP endpoints and framework-based APIs, as well as secrets, storage, observability, and failure-aware operations. The emphasis is on architectural judgment: when to choose denser container utilization over more replicas, how to avoid deployment-time packaging failures, and how to design systems that remain correct under retries, cold starts, and scaling pressure.

The book assumes strong Python fluency and comfort with modern software delivery and cloud concepts. Its distinguishing strength is depth: each topic is framed operationally, with close attention to version-sensitive guidance, trade-offs, and production realities that matter once prototypes b

Tietoa kirjasta

"Serverless AI and Data with Modal: Running GPU Workloads, Pipelines, and Apps Without Managing Infrastructure"

For experienced Python engineers, ML practitioners, and platform-minded developers, this book offers a rigorous guide to building serious AI and data systems on Modal without operating clusters, schedulers, or container fleets. Rather than treating serverless as a convenience feature, it presents Modal as an execution model for production workloads—one that changes how you think about deployment boundaries, GPU access, autoscaling, and the relationship between code, runtime, and infrastructure.

Across the book, readers learn how to package reproducible runtime environments, deploy remote functions and persistent applications, tune concurrency and warm capacity, and run GPU-backed inference, training, and batch jobs with predictable performance. It also covers map-style distributed pipelines, scheduled automation, HTTP endpoints and framework-based APIs, as well as secrets, storage, observability, and failure-aware operations. The emphasis is on architectural judgment: when to choose denser container utilization over more replicas, how to avoid deployment-time packaging failures, and how to design systems that remain correct under retries, cold starts, and scaling pressure.

The book assumes strong Python fluency and comfort with modern software delivery and cloud concepts. Its distinguishing strength is depth: each topic is framed operationally, with close attention to version-sensitive guidance, trade-offs, and production realities that matter once prototypes b

Aloita kirja saman tien hintaan 0 €

  • Kokeilujakson aikana käytössäsi on kaikki sovelluksen kirjat
  • Ei sitoumusta, voit perua milloin vain
Kokeile nyt ilmaiseksi
Yli 52 000 ihmistä on antanut Nextorylle viisi tähteä App Storessa ja Google Playssä.