Skip to content

Repository files navigation

xpuruntime

A unified GPU/NPU execution runtime that bridges research and production by making engine and kernel selection explicit, reproducible, and framework-agnostic.

"연구에서 검증한 실행 전략(커널/정밀도/메모리 정책)을 그대로 서비스까지 가져간다."

Overview

xpuruntime is not a new model framework—it is a runtime library that controls GPU/NPU execution. It provides:

  • Policy as Code: Kernel/engine selection as explicit, versioned policy
  • Unified policy model: Same policy concepts across training and inference
  • Framework-agnostic: Reusable tuning assets across PyTorch, TensorFlow, ONNX Runtime, TensorRT
  • Explainable execution: Logging and profiling of which kernels/engines/precision were used
Layer Role Language
Control Plane User UX, API, pipeline definition Python
Data Plane Execution runtime (memory, stream, kernel, engine) C++
Performance Layer Engine/kernel internals CUDA etc.

Install

Default install is pure Python (no CUDA/CMake required):

pip install -e ".[dev]"   # editable + dev deps

The C++/CUDA extension is optional; when CUDA and CMake are available, see 07_build_packaging.md for native build.

Optional extras (e.g. Intel GPU device count and inference):

pip install xpuruntime[intel]      # Intel GPU count via OpenCL (pyopencl)
pip install xpuruntime[openvino]   # Intel GPU/CPU inference; no manual OpenVINO setup

Example: run inference on Intel GPU (device selection is handled for you):

from xpuruntime.inference import create_openvino_session
import numpy as np

session = create_openvino_session("model.onnx", device="auto")  # auto = Intel GPU if available
outputs = session.run({session.input_names[0]: your_input_array})

See examples/intel_gpu_openvino_demo.py.

Documentation

문서 설명
00_overview 비전, 타깃 사용자, 차별화
01_architecture 아키텍처
02_cpp_core_runtime C++ 코어
03_python_binding Python 바인딩
04~08 Inference, Training, Kernel policy, Build, Testing
tasks/ TASK_001~014, task_log

Open Source & Contributing

xpuruntime is open source and maintained as a nonprofit, community-driven project. We welcome contributions: bug reports, feature ideas, docs, and code. Issues tagged with good first issue or help wanted are a good place to start.

  • CONTRIBUTING.md – 개발 환경 설정, PR 절차, 코드 스타일
  • Issues – 버그 리포트, 기능 제안, 기여 이슈

License

This project is licensed under the Apache License 2.0. See LICENSE for the full text.

About

Unified GPU/NPU execution runtime for ML. Open source, community-driven.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages