Magia en cada página | Especial Harry Potter hasta 55% OFF  Ver más

Enviar a
C.A.B.A., Ciudad Autónoma de Buenos Aires
0
  • argentina
  • chile
  • colombia
  • españa
  • méxico
  • perú
  • estados unidos
  • internacional

Selecciona tu país

América

Europa

Resto del mundo

portada Parallel Computing for AI and ML Engineers. Build Scalable Deep Learning Systems with GPU Programming, Multi-GPU Training, and Production Workloads (en Inglés)
Formato
Libro Físico
Año
2026
Idioma
Inglés
N° páginas
436
Encuadernación
Tapa Blanda
Dimensiones
28.00 x 21.60 x 2.20 cm
ISBN13
9798195370404

Parallel Computing for AI and ML Engineers. Build Scalable Deep Learning Systems with GPU Programming, Multi-GPU Training, and Production Workloads (en Inglés)

M.t Holbrook (Autor) · Independently published · Tapa Blanda

Parallel Computing for AI and ML Engineers. Build Scalable Deep Learning Systems with GPU Programming, Multi-GPU Training, and Production Workloads (en Inglés) - M.T Holbrook

Libro Nuevo Importado
Envío: 15 a 21 días háb.
$ 159.090$ 79.545
-50%
Costos de importación incluídos en el precio ✅
Libro Nuevo

Quedan 100 unidades

$ 79.545
Llega entre el 24 Ago y el 01 Sep a C.A.B.A., Ciudad Autónoma de Buenos Aires. Seleccionar ubicación

Reseña del libro "Parallel Computing for AI and ML Engineers. Build Scalable Deep Learning Systems with GPU Programming, Multi-GPU Training, and Production Workloads (en Inglés)"

Stop Guessing. Start Building ML Systems That Actually Scale.

Most ML engineers learn GPU computing the hard way - through production failures, mysterious hangs, and models that take three times longer to train than they should. This book gives you the understanding and the tools to get it right the first time.

What This Book Covers

-GPU architecture internals: CUDA cores, warps, shared memory, and memory coalescing

-Writing and optimizing custom CUDA kernels in C++

-Data parallel, model parallel, and pipeline parallel training with PyTorch DDP and FSDP

-Multi-node training with NCCL, MPI, and InfiniBand

-Mixed precision training and gradient scaling

-ZeRO optimizer stages 1, 2, and 3 with DeepSpeed

-Custom DataLoader optimization and NVIDIA DALI

-Production model serving with Triton Inference Server

-Kubernetes deployment with GPU autoscaling

-Complete profiling workflows with Nsight and PyTorch Profiler

-Troubleshooting CUDA OOM, NCCL hangs, and NaN losses

-Capacity planning and hardware selection for real workloads

Who This Book Is For

This book is written for ML engineers, AI researchers, and software engineers working on deep learning infrastructure who want to move beyond single-GPU experiments and build systems that perform at scale. You should be comfortable with Python and have basic familiarity with PyTorch or TensorFlow. No prior CUDA experience required.

What Makes This Book Different

Every chapter includes complete, runnable code. Architecture diagrams show how components connect. Benchmark results come from real hardware measurements. The troubleshooting appendices address the exact errors that stop real training jobs. This is not a survey of techniques. It is a working engineer's guide to building production parallel ML systems.

Opiniones del libro

Preguntas frecuentes sobre el libro

Todos los libros de nuestro catálogo son Originales.
El libro está escrito en Inglés.
La encuadernación de esta edición es Tapa Blanda.

Preguntas y respuestas sobre el libro

¿Tienes una pregunta sobre el libro? Inicia sesión para poder agregar tu propia pregunta.

Opiniones sobre Buscalibre

Ver más opiniones de clientes