Ollama: Self-Hosted Large Language Models Replacing OpenAI and Claude APIs Training Course
Ollama is an open-source tool for running large language models locally on consumer and enterprise hardware. It abstracts model quantization, GPU allocation, and API serving into a single command-line interface, enabling organizations to self-host LLMs like Llama, Mistral, and Qwen without sending prompts or data to OpenAI, Anthropic, or Google.
This instructor-led, live training (online or onsite) is aimed at intermediate AI engineers and platform operators who wish to use Ollama to replace cloud LLM APIs with self-hosted, sovereign language model inference.
By the end of this training, participants will be able to:
- Install Ollama on Linux, macOS, and Windows with GPU support.
- Pull, quantize, and serve models from the Ollama registry and HuggingFace.
- Build custom Modelfiles with system prompts and parameter tuning.
- Integrate local LLMs with applications via the OpenAI-compatible API.
- Optimize inference performance for CPU-only and multi-GPU setups.
Format of the Course
- Interactive lecture and discussion.
- Lots of exercises and practice.
- Hands-on implementation in a live-lab environment.
Course Customization Options
- To request a customized training for this course, please contact us to arrange.
Course Outline
AI Sovereignty and LLM Local Deployment
- Risks of cloud LLMs: data retention, training on inputs, foreign jurisdiction.
- Ollama architecture: model server, registry, and OpenAI-compatible API.
- Comparison with vLLM, llama.cpp, and Text Generation Inference.
- Model licensing: Llama, Mistral, Qwen, and Gemma terms.
Installation and Hardware Setup
- Installing Ollama on Linux with CUDA and ROCm support.
- CPU-only fallback and AVX/AVX2 optimization.
- Docker deployment and persistent volume mapping.
- Multi-GPU setup and VRAM allocation strategies.
Model Management
- Pulling models from the Ollama registry: ollama pull llama3.
- Importing GGUF models from HuggingFace and TheBloke.
- Quantization levels: Q4_K_M, Q5_K_M, Q8_0 tradeoffs.
- Model switching and concurrent model loading limits.
Custom Modelfiles
- Writing Modelfile syntax: FROM, PARAMETER, SYSTEM, TEMPLATE.
- Temperature, top_p, and repeat_penalty tuning.
- System prompt engineering for role-specific behavior.
- Creating and publishing custom models to local registry.
API Integration
- OpenAI-compatible /v1/chat/completions endpoint.
- Streaming responses and JSON mode.
- Integrating with LangChain, LlamaIndex, and custom apps.
- Authentication and rate limiting with reverse proxy.
Performance Optimization
- Context window sizing and KV cache management.
- Batch inference and parallel request handling.
- CPU thread allocation and NUMA awareness.
- Monitoring GPU utilization and memory pressure.
Security and Compliance
- Network isolation for model serving endpoints.
- Input filtering and output moderation pipelines.
- Audit logging of prompts and completions.
- Model provenance and hash verification.
Requirements
- Intermediate Linux and container administration.
- Understanding of machine learning and transformer models at high level.
- Familiarity with REST APIs and JSON.
Audience
- AI engineers and developers replacing cloud LLM APIs.
- Organizations with data sensitivity preventing cloud model usage.
- Government and defense teams requiring air-gapped language models.
Open Training Courses require 5+ participants.
Ollama: Self-Hosted Large Language Models Replacing OpenAI and Claude APIs Training Course - Booking
Ollama: Self-Hosted Large Language Models Replacing OpenAI and Claude APIs Training Course - Enquiry
Ollama: Self-Hosted Large Language Models Replacing OpenAI and Claude APIs - Consultancy Enquiry
Upcoming Courses
Related Courses
Advanced Ollama Model Debugging & Evaluation
35 HoursThe Advanced Ollama Model Debugging & Evaluation course provides a comprehensive deep dive into diagnosing, testing, and assessing model behavior within local or private Ollama deployments.
Delivered as live, instructor-led training available either online or on-site, this program is tailored for advanced AI engineers, MLOps professionals, and QA practitioners who aim to guarantee the reliability, accuracy, and operational readiness of Ollama-based models in production environments.
Upon completion of this training, participants will gain the ability to:
- Conduct systematic debugging of Ollama-hosted models and reliably reproduce failure scenarios.
- Design and implement robust evaluation pipelines incorporating both quantitative and qualitative metrics.
- Establish observability practices (including logs, traces, and metrics) to monitor model health and detect drift.
- Automate testing, validation, and regression checks integrated into CI/CD pipelines.
Course Format
- Interactive lectures and discussions.
- Practical labs and debugging exercises utilizing Ollama deployments.
- Case studies, group troubleshooting sessions, and automation workshops.
Customization Options
- To request a customized version of this course, please contact us to arrange it.
Building Private AI Workflows with Ollama
14 HoursThis instructor-led, live training in Czech Republic (online or onsite) is designed for advanced professionals who aim to implement secure and efficient AI-driven workflows using Ollama.
By the end of this training, participants will be able to:
- Deploy and configure Ollama for private AI processing.
- Integrate AI models into secure enterprise workflows.
- Optimize AI performance while maintaining data privacy.
- Automate business processes with on-premise AI capabilities.
- Ensure compliance with enterprise security and governance policies.
Deploying and Optimizing LLMs with Ollama
14 HoursThis instructor-led live training in Czech Republic (online or onsite) is designed for intermediate professionals aiming to deploy, optimize, and integrate LLMs using Ollama.
By the end of this training, participants will be able to:
- Set up and deploy LLMs using Ollama.
- Optimize AI models for performance and efficiency.
- Leverage GPU acceleration for improved inference speeds.
- Integrate Ollama into workflows and applications.
- Monitor and maintain AI model performance over time.
EXO: End-to-End Local AI Cluster Deployment
21 HoursThis instructor-led, live training in Czech Republic (online or onsite) is aimed at system administrators and DevOps engineers who wish to deploy, configure, and manage EXO clusters for private LLM inference across multiple Apple Silicon or Linux nodes.
EXO for DevOps: Building Private AI Infrastructure
21 HoursThis instructor-led, live training in Czech Republic (online or onsite) is aimed at DevOps engineers and infrastructure architects who wish to automate the provisioning, monitoring, and lifecycle management of private AI clusters built on EXO.
EXO Security and Governance: Offline Model Management
14 HoursThis instructor-led, live training in Czech Republic (online or onsite) is aimed at security engineers and compliance officers who wish to harden EXO deployments, control model access, and govern AI workloads running entirely on-premise.
Fine-Tuning and Customizing AI Models on Ollama
14 HoursThis instructor-led live training in Czech Republic (online or on-site) is designed for advanced professionals who wish to fine-tune and customize AI models on Ollama to improve performance and support domain-specific applications.
By the end of this training, participants will be able to:
- Set up an efficient environment for fine-tuning AI models on Ollama.
- Prepare datasets for supervised fine-tuning and reinforcement learning.
- Optimize AI models for performance, accuracy, and efficiency.
- Deploy customized models in production environments.
- Evaluate model improvements and ensure robustness.
Secure Local Agentic AI: On-Prem Ollama Development for Regulated Industries
21 HoursThis instructor-led, live training in Czech Republic (online or onsite) is designed for developers and technical teams who wish to use Ollama and open models to build and run private agentic AI solutions on internal infrastructure.
By the end of this training, participants will be able to install and configure Ollama, evaluate and run open models locally, create simple agentic and retrieval-based workflows, and apply security and governance controls for regulated environments.
Multimodal Applications with Ollama
21 HoursOllama serves as a powerful platform for executing and fine-tuning large language models as well as multimodal models directly on local infrastructure.
This instructor-led training session, available either online or at your location, is specifically designed for advanced machine learning engineers, AI researchers, and product developers aiming to construct and deploy sophisticated multimodal applications using Ollama.
Upon completing this training, participants will be equipped to:
- Configure and execute multimodal models utilizing Ollama.
- Seamlessly combine text, image, and audio inputs for practical, real-world solutions.
- Create systems for document comprehension and visual question answering.
- Develop intelligent multimodal agents capable of cross-modal reasoning.
Course Delivery Format
- Engaging lectures paired with interactive discussion.
- Practical exercises using authentic multimodal datasets.
- Live laboratory sessions focusing on implementing multimodal pipelines via Ollama.
Customization Opportunities
- For those seeking a tailored training experience, please reach out to us to arrange a customized schedule.
Getting Started with Ollama: Running Local AI Models
7 HoursThis instructor-led, live training in Czech Republic (online or onsite) is aimed at beginner-level professionals who wish to install, configure, and use Ollama for running AI models on their local machines.
By the end of this training, participants will be able to:
- Understand the fundamentals of Ollama and its capabilities.
- Set up Ollama for running local AI models.
- Deploy and interact with LLMs using Ollama.
- Optimize performance and resource usage for AI workloads.
- Explore use cases for local AI deployment in various industries.
Ollama & Data Privacy: Secure Deployment Patterns
14 HoursOllama is a platform designed for running large language and multimodal models locally, while also supporting robust secure deployment strategies.
This instructor-led, live training (available online or onsite) targets intermediate-level professionals looking to deploy Ollama with stringent data privacy and regulatory compliance measures.
Upon completion of this training, participants will be capable of:
- Deploying Ollama securely within containerized and on-premises environments.
- Applying differential privacy techniques to protect sensitive data.
- Implementing secure logging, monitoring, and auditing practices.
- Enforcing data access controls in alignment with compliance requirements.
Course Format
- Interactive lectures and discussions.
- Hands-on labs focused on secure deployment patterns.
- Compliance-oriented case studies and practical exercises.
Course Customization Options
- To request a customized version of this course, please contact us to arrange.
Ollama Applications in Finance
14 HoursOllama is a lightweight platform for running large language models locally.
This instructor-led, live training (online or onsite) is aimed at intermediate-level finance practitioners and IT personnel who wish to implement, customize, and operationalize Ollama-based AI solutions in financial environments.
By completing this training, participants will gain the skills needed to:
- Deploy and configure Ollama for secure use in financial operations.
- Integrate local LLMs into analytical and reporting workflows.
- Adapt models to finance-specific terminology and tasks.
- Apply security, privacy, and compliance best practices.
Format of the Course
- Interactive lecture and discussion.
- Hands-on financial data exercises.
- Live-lab implementation of finance-focused scenarios.
Course Customization Options
- To request a customized training for this course, please contact us to arrange.
Ollama Applications in Healthcare
14 HoursOllama is a lightweight platform for running large language models locally.
This instructor-led, live training (online or onsite) is aimed at intermediate-level healthcare practitioners and IT teams who wish to deploy, customize, and operationalize Ollama-based AI solutions within clinical and administrative environments.
Upon completing this training, participants will be able to:
- Install and configure Ollama for secure use in healthcare settings.
- Integrate local LLMs into clinical workflows and administrative processes.
- Customize models for healthcare-specific terminology and tasks.
- Apply best practices for privacy, security, and regulatory compliance.
Format of the Course
- Interactive lecture and discussion.
- Hands-on demonstrations and guided exercises.
- Practical implementation in a sandboxed healthcare simulation environment.
Course Customization Options
- To request a customized training for this course, please contact us to arrange.
Ollama for Responsible AI and Governance
14 HoursOllama is a platform designed for locally running large language and multimodal models, with a focus on supporting governance and responsible AI practices.
This instructor-led live training (available online or onsite) is targeted at intermediate to advanced-level professionals who want to embed fairness, transparency, and accountability into applications powered by Ollama.
Upon completing this training, participants will be capable of:
- Applying responsible AI principles during Ollama deployments.
- Implementing strategies for content filtering and bias mitigation.
- Designing governance workflows that ensure AI alignment and auditability.
- Establishing monitoring and reporting frameworks to meet compliance requirements.
Course Format
- Interactive lectures and discussions.
- Hands-on labs focused on designing governance workflows.
- Case studies and exercises centered on compliance.
Customization Options
- To arrange a customized version of this course, please contact us.
Sovereign AI for Regulated Organizations: Controlling Data, Models and Inference Environments
7 HoursThis instructor-led, live training in Czech Republic (online or onsite) is aimed at intermediate-level IT leaders, compliance professionals, security teams, and enterprise architects who wish to use sovereign AI principles and governance practices to design AI environments that protect sensitive data, support localization requirements, and reduce vendor lock-in.
By the end of this training, participants will be able to: explain sovereign AI concepts, evaluate hosting and governance options, define controls for prompts and logs, and create a practical adoption roadmap.