Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Core Performance Concepts and Metrics
- Analyzing latency, throughput, power consumption, and resource utilization
- Distinguishing between system-wide and model-specific bottlenecks
- Differentiating profiling approaches for inference versus training
Profiling Techniques on Huawei Ascend
- Utilizing CANN Profiler and MindInsight
- Diagnosing kernel and operator performance
- Analyzing offload patterns and memory mapping
Performance Analysis on Biren GPU
- Leveraging Biren SDK performance monitoring capabilities
- Exploring kernel fusion, memory alignment, and execution queues
- Implementing power and temperature-aware profiling
Profiling on Cambricon MLU
- Using BANGPy and Neuware performance utilities
- Interpreting kernel-level logs for visibility
- Integrating the MLU profiler with deployment frameworks
Graph and Model-Level Optimization Strategies
- Implementing graph pruning and quantization techniques
- Applying operator fusion and computational graph restructuring
- Standardizing input sizes and tuning batch configurations
Memory and Kernel Efficiency
- Refining memory layout and data reuse patterns
- Managing buffers efficiently across different chipsets
- Applying platform-specific kernel-level tuning methods
Best Practices for Cross-Platform Optimization
- Achieving performance portability through abstraction strategies
- Developing shared tuning pipelines for multi-chip environments
- Case study: Tuning an object detection model across Ascend, Biren, and MLU
Conclusion and Recommended Next Steps
Requirements
- Practical experience with AI model training or deployment pipelines
- Sound understanding of GPU/MLU compute principles and model optimization strategies
- Fundamental knowledge of performance profiling tools and key metrics
Target Audience
- Performance engineers
- Machine learning infrastructure teams
- AI system architects
21 Hours