Altos aiWorks

Altos aiWorks Artificial Intelligence LLM & AI Developer GPU Resource Management Platform

Altos aiWorks

Powering the Enterprise AI Factory of Tomorrow

As enterprises build out their own AI Factory, GPU resources are no longer just a hardware investment — they demand intelligent scheduling, allocation, and multi-tenant management to truly deliver on that investment's potential. Altos aiWorks is the essential GPU resource management and scheduling platform for enterprises building their AI Factory, purpose-built for LLM users, AI developers, and resource administrators alike.
Combining exceptional flexibility with intelligent CPU, memory, and GPU allocation, aiWorks is built on a Kubernetes container cluster architecture that streamlines LLM development and inference service deployment — accelerating AI application rollout while maximizing resource efficiency.
Altos aiWorks 5.0 features an industry-leading graphical NVIDIA Dynamo workspace, paired with integrated Ceph high-availability storage — giving users a more stable, more elastic foundation for faster development cycles and stronger economic returns, driving AI innovation forward into the LLM era.

Enterprise Identity & Access Management

NEW Altos aiWorks provides a three-tier role-based permission architecture that integrates seamlessly with existing enterprise LDAP/AD account systems. Login supports OTP two-factor authentication, which enterprises can enforce as a mandatory requirement in line with their security policies.

Enterprise Identity & Access Management

Flexible Storage & Resource Management

Supports Three Storage Types
• Private Disk: Native Persistent Volume (PVC) with storage quota management.
• NFS: Seamlessly integrates with existing network storage for fast shared data access.
• Samba: Supports cross-platform access, including Windows, with storage quotas and centralized management.
Altos aiWorks also integrates NVIDIA GPUDirect Storage (GDS) to reduce CPU overhead and storage access latency, accelerating data processing for large AI model training, inference, and HPC workloads.

Flexible Storage & Resource Management

Accelerated AI Development & Inference

Altos aiWorks 5.0 comes with built-in support for popular development environments — JupyterLab, PyTorch, TensorFlow, and Open WebUI — which users can launch with a single click for rapid model development, training, and testing.
NEW It also features an industry-leading graphical NVIDIA Dynamo inference workspace, allowing users to deploy and configure model inference services entirely through the interface, with no command-line operations required.

Accelerated AI Development & Inference

Intelligent Workspace Management

Altos aiWorks 5.0 supports web-based Workspace management, letting users create, monitor, and manage the full lifecycle of Workspaces directly through the browser — no command-line operations needed. Resource release schedules can also be configured at the Group level, automatically freeing up resources once idle or expired, preventing prolonged resource lock-up by a single user or group and improving overall resource efficiency.

Intelligent Workspace Management

Usage Monitoring & Cost Management

NEWBilling in Altos aiWorks 5.0 spans three core capabilities: pricing configuration by CPU, RAM, storage, and GPU type; billing management, with detailed deduction records, date filtering, and data export; and billing analytics, presenting cost trends through charts and itemized breakdowns of CPU, GPU, RAM, and storage expenses.

Usage Monitoring & Cost Management

One-Stop Integration

Altos aiWorks 5.0 integrates deeply with Altos BrainSphere™ servers, delivering seamless continuity from the hardware layer to the AI application layer. Enterprises no longer need to stitch together multiple systems — everything from hardware deployment to AI service launch is available in a single, true one-stop AI Factory experience.

One-Stop Integration

Supports NVIDIA Multi-Instance GPU Technology

Altos aiWorks is an industry-leading AI computing platform that supports NVIDIA A100 Multi-Instance GPU (MIG) technology. By partitioning GPUs—ensuring isolated high-bandwidth memory, cache, and compute cores—it seamlessly handles workloads of any scale, accelerates resource scalability, and maximizes overall utilization.

Multi-Instance GPU (MIG)

Multi-Instance GPU (MIG)

Physical partitioning that splits a GPU into multiple instances, ensuring strict isolation and guaranteed QoS for each.

Multi-Process Service (MPS)

Multi-Process Service (MPS)

A tool that enables compute kernels submitted by multiple CPU processes to execute concurrently on the same GPU.

Dedicated GPU Allocation

Dedicated GPU Allocation

Allocate more than one GPU.

Altos aiWorks v5.0 Platform Architecture Overview

Built on Kubernetes, the platform unifies GPU, networking, storage, container images, resource scheduling, user management, and system monitoring for centralized AI infrastructure management.

Empower Organizations to Deploy AI with Maximum Efficiency

Altos aiWorks integrates AI development, deployment, resource management, and inference monitoring into a unified platform. Built for enterprises and academia, it simplifies infrastructure management to accelerate the landing of Generative AI and Agentic AI workloads, driving ultimate efficiency in AI adoption.

Recommended with Altos BrainSphere™ Workstations / Servers

Altos aiWorks solution offers a wide range of AI systems from server to workstation, supporting multiple GPU AI accelerators on demand to meet your current and growing needs.