IP Portfolio

Complete GPU IP Capability System Glenfly has built a complete IP portfolio spanning 3D Graphics, 2D, GPGPU, matrix computing, video encoding/decoding and post-processing, display processing (PHY & Controller), Network-on-Chip (NoC), and data compression. It provides core capabilities for graphics rendering, general-purpose computing, AI acceleration, video processing, display output, and SoC integration.

Flexible integration for different systems Each IP can be deployed independently, or flexibly combined and integrated into a unified solution according to chip architecture, product form and application scenario, helping customers achieve efficient adaptation, rapid integration and differentiated design across different chip platforms.

Architecture-level collaborative optimization Built on an in-house, controllable underlying hardware architecture together with system scheduling and resource-management capabilities, Glenfly unifies data formats, layout specifications and communication protocols within its IP system and integrates consistent compression mechanisms and security policies. Through architecture-level collaborative optimization, data flow, communication paths and resource scheduling among IPs are efficiently connected, continuously improving bandwidth utilization and data-exchange efficiency at the system level while reducing runtime overhead and integration complexity.

Leveraging continuously evolving software-hardware co-design, Glenfly provides customers with customized IP and scalable chip-level solutions, helping products achieve a better balance across key metrics such as performance, power and area.

Configurable GPGPU & Matrix

General-purpose Computing · Matrix Computing · High-bandwidth Memory · AI Ecosystem

A configurable GPGPU IP platform for general-purpose computing and AI inference/training, integrating a SIMT architecture, Matrix Core and Matrix Fabric Memory to deliver powerful general-purpose compute capabilities, flexible configuration and a complete software ecosystem, helping customers rapidly build high-performance computing platforms.

SIMT architecture

Optional SFU / VLIW / timely multi-issue and other architectural features

Matrix Core

Fully configurable, supporting multiple data formats

Matrix Fabric Memory

High-bandwidth matrix data movement and accumulator-buffer optimization

Energy-efficient design

Locally expandable, background-resident power-saving design

Applications

HD display devices
Multimedia processing chips
video output & receiving scenarios
cloud gaming · cloud desktop · conferencing systems

Glenfly HUCA software-stack ecosystem

Supports mainstream frameworks and model formats such as ONNX, Llama, PyTorch, Caffe and TensorFlow

Drivers AI operator library Model-conversion toolchain Compiler Deep learning framework adaptation Deployment management platform Runtime LLM training & inference engine

Polyhedral Model

Find a linear "time" function

T(i,j) = c0 + c1j

with c · d > 0 for every d ∈ D.

All dependences go "forward in time".

Parallelism = all points on the same hyperplane.

Wavefronts advance with increasing k.

Matrix Fabric Memory

(Configurable)

Configurable capacity, optimized for Matrix Core accumulator buffering and high-bandwidth matrix data movement.

Compute-core configuration

FP32 ALU core count

1 Core = 2 FLOPs /cycle

Matrix Core performance

INT8 OPS /cycle

Cache configuration

(Configurable)

1/2 MB 1 MB 2 MB 4 MB+ 8 MB+

Precision and
data-format support

(Configurable)

INT8 INT16 FP4 FP8 FP16 BF16 TF32