Get in Touch

Course Outline

Introduction to Biren GPU Architecture

  • Overview of Biren and its key use cases.
  • Hardware configuration: cores, memory structures, and compute clusters.
  • Comparative analysis with NVIDIA and AMD GPUs.

Configuring the Biren Programming Environment

  • Installing the Biren SDK and runtime components.
  • Understanding the toolchain and compiler architecture.
  • Essential project structures and build workflows.

GPU Programming with the Biren Stack

  • Thread and block organization models.
  • Memory management strategies and data transfer mechanisms.
  • Kernel development and launch patterns.

Migrating from CUDA to Biren

  • Techniques for translating CUDA code.
  • Common API mappings and necessary adaptations.
  • Practical labs focused on code conversion.

Debugging and Profiling

  • Utilizing Biren’s debugger and profiler tools.
  • Identifying system bottlenecks.
  • Optimizing memory access patterns.

Optimization Techniques

  • Thread scheduling and instruction pipelining.
  • Loop unrolling and efficient shared memory usage.
  • Advanced kernel tuning for maximum throughput.

Case Studies and Application Examples

  • Training models using Biren accelerators.
  • Porting and profiling vision or NLP models.
  • Performance comparison against CUDA/NVIDIA environments.

Summary and Next Steps

Requirements

  • Fundamental knowledge of GPU architecture and parallel processing concepts.
  • Hands-on experience with CUDA, OpenCL, or comparable GPU programming frameworks.
  • Proficiency with deep learning frameworks such as PyTorch or TensorFlow.

Target Audience

  • HPC developers.
  • AI infrastructure engineers.
  • Specialists in performance optimization.
 21 Hours

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories