Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Overview of the Chinese AI GPU Ecosystem
- Comparison of Huawei Ascend, Biren, and Cambricon MLU
- Contrast between CUDA and CANN, Biren SDK, and BANGPy models
- Industry trends and vendor ecosystems
Preparing for Migration
- Evaluating your existing CUDA codebase
- Selecting target platforms and SDK versions
- Installing toolchains and configuring the environment
Code Translation Techniques
- Migrating CUDA memory access patterns and kernel logic
- Mapping compute grid and thread models
- Automated versus manual translation methods
Platform-Specific Implementations
- Utilizing Huawei CANN operators and custom kernels
- The Biren SDK conversion pipeline
- Reconstructing models using BANGPy (Cambricon)
Cross-Platform Testing and Optimization
- Profiling execution on each target platform
- Comparing memory tuning and parallel execution
- Monitoring performance and iterative refinement
Managing Mixed GPU Environments
- Implementing hybrid deployments with multiple architectures
- Developing fallback strategies and device detection
- Using abstraction layers to enhance code maintainability
Case Studies and Best Practices
- Migrating vision and NLP models to Ascend or Cambricon
- Adapting inference pipelines for Biren clusters
- Resolving version mismatches and API discrepancies
Summary and Next Steps
Requirements
- Proficiency in programming with CUDA or GPU-based applications
- Knowledge of GPU memory models and compute kernels
- Familiarity with AI model deployment or acceleration workflows
Audience
- GPU developers
- System architects
- Porting engineers
21 Hours