PreMize
Sign in
NVIDIA DGX SPARK · GB10 · 128 GB UNIFIED

Efficient Local LLMon one DGX Spark.

Upload a document, ask in plain language, and get answers that cite their exact page — from an AI that runs entirely on one desk-side machine. Your documents never leave it.

BETA · NO SIGN-UP · 15 MIN FREE
0.66 s
to first token
16
concurrent chats
125 tok/s
sustained
Measured on the production DGX Spark
IN PARTNERSHIP WITH
01THE LAYERS

Every layer, engineered by experts.

What feels effortless on the surface is a full AI infrastructure underneath — memory budgets, model serving, retrieval, governance, and experience, designed together as one system on one machine.

Effortless to use. Efficient by design.

Every layer beneath is researched, tuned, and optimized — to deliver maximum performance for you.

02THE MACHINE

One machine runs everything.

No rack, no cloud, no integration project. The whole platform — models, index, and dashboards — lives in one quiet box on a desk.

  • GB10 GRACE-BLACKWELL
  • 128 GB UNIFIED MEMORY
  • 1 MACHINE — EVERYTHING
NVIDIA DGX Spark
NVIDIA GB10 Grace-Blackwell superchip board beside the DGX Spark chassis

The GB10 superchip: CPU and GPU on one die, sharing 128 GB of unified memory.

03THE MODELS

Intelligence, sized to the machine.

Two chat models and a Thai document reader share one memory pool — the platform decides what is awake, so users never think about it.

PreMize v0.1

35B · 262,144-TOKEN CONTEXT

The primary model for chat and summarization — deep answers over long documents.

PreMize Lite

8B · WAKES ON DEMAND

A fast secondary model that sleeps when idle and wakes automatically when called.

Thai OCR

TH + EN · SCANNED DOCUMENTS

Reads scanned Thai paperwork page by page — contracts, handbooks, invoices.

Semantic search

MEANING, NOT JUST KEYWORDS

Finds the right passage even when the question uses different words than the document.

04HOW IT WORKS

Upload. Ask. Get cited answers.

No setup, no training, no manual. The platform reads, summarizes, and indexes every document automatically — then answers in plain language with the page number attached.

01UPLOADDrop in a PDF — born-digital or scanned.02UNDERSTANDRead, summarized, and indexed automatically.03ASK — GET CITED ANSWERSEvery answer carries its source page.
05MEASURED PERFORMANCE

Measured, not promised.

0.00 s
time to first token
at 16 concurrent chats
0×
faster first token under load
after expert tuning
0 tok/s
aggregate throughput
sustained at saturation
0
errors
/ 192-request soak

Recorded on the production DGX Spark under concurrent load — not a projection.

06PROVENANCE

Every answer cites its page.

Citations come from the document itself — deterministic “Ref page N” links to the exact source page. The model can never invent one. Trust is checkable in one click.

How many days of annual leave do employees receive?

Full-time employees receive 12 days of annual leave per year.

Ref page 12
handbook_th.pdf — PAGE 12
07THE PLATFORM

Complete, and effortless.

Cited document chat

Ask across one document or your whole library — answers carry page references.

Built-in assistant

A floating guide that knows the platform and helps every user find their way.

Agent Studio

Compose agents visually; a test scorecard must pass before anything goes live.

Customer channels

Public web chat and LINE OA, answering customers with the same cited accuracy.

Open API

Standard-compatible API keys, self-served — connect your own tools in minutes.

Usage dashboards

Per-user metering and live system health, visible at a glance.

08THE PRODUCT

The product, as it ships.

Real screens from the running platform — Thai-first, cited, and observable end to end.

Cited document chat

Cited document chat

Summary, conversation, and citations side by side — every claim carries a “Ref page N” chip that links straight to its source page.

Agent Studio

Agent Studio

Compose an agent from tools, document scope, and policies on a visual canvas — and it must pass a test scorecard before it goes live.

Admin control tower

Admin control tower

Tokens, requests, and cost — per user, per feature, per model. Fourteen days of the whole organization at a glance.

Chat with live web search

Chat with live web search

An optional governed web search adds fresh sources to general chat, each answer carrying its reference chips.

Model picker

Model picker

Switch between the 35B and the fast 8B — capabilities shown up front.

Live GPU telemetry

Live GPU telemetry

Utilization, temperature, and power of the GB10, streamed in real time.

Self-serve API keys

Self-serve API keys

A standard-compatible endpoint with per-user metering, one click away.

09HERMES AGENT
Hermes Agent

Your machine, in your pocket.

Hermes is an intelligent agent that lives on the DGX Spark and answers on Telegram. Wake models, check the GPU, ask about your documents — from your phone, anywhere, anytime. The conversation travels; the computing never leaves the machine.

  • Chat on Telegram from any phone — anywhere, anytime
  • Wake and sleep models, run prompts, check live GPU status
  • Everything is processed locally on your DGX Spark

Powered by Hermes · Nous Research

TELEGRAM
Wake the fast model and give me a machine status.
Done — PreMize Lite is awake ✓ GPU 5% · 44°C · 12 W. Ready when you are.
A real Telegram session: the agent identifies a photo and answers with live web results
REAL SESSION · @PREMIZE_BOT · PHOTO + LIVE WEB SEARCH
SELECTED TOOLS · ALL ON DGX SPARK

A hand-picked toolset — each tool tuned and measured to run at its best on the appliance, and every action passes policy → approval → audit.

system_statusREAD

GPU load, temperature, power, and service health — live from the machine.

model_controlGOVERNED

Wake and sleep models from chat; sensitive actions wait for your approval.

doc_searchREAD

Ask about your documents — answers cite their exact page.

web_searchNETWORK

Live web results through the appliance's own search engine — no third-party API.

visionLOCAL

Send a photo in the chat; the machine reads it and explains what it sees.

usage_approvalsREAD

Token usage by model, plus the approvals inbox — decide right in the chat.

10WHY PREMIZE

The same machine. A different class of system.

A bare DGX Spark gives you excellent hardware — and months of hard engineering before it serves your business. PreMize ships with that engineering already done: we spent the research hours on memory budgets, serving behavior, retrieval, and governance so every layer runs at its optimum — resources, latency, and efficiency. Complex underneath, so it is simple on top.

DGX Spark only

DIY

excellent hardware — everything else is homework

  • Partitioning 128 GB of unified memory (trial, error, OOM)
  • Model serving, quantization, and tuning
  • Thai OCR pipeline for scanned documents
  • RAG with trustworthy page-level citations
  • Governance, approvals, and audit
  • Dashboards, metering, and monitoring

DGX Spark + PreMize

READY

engineered, measured, ready on day one

  • Every gigabyte of unified memory budgeted
  • Serving tuned for the GB10 — with an auto sleep/wake steward
  • Thai OCR built in, page by page
  • Cited answers out of the box — “Ref page N”
  • Policy → approval → audit, from day one
  • Live dashboards, per-user metering, open API

Measured on the same DGX Spark — before and after our tuning

Time to first token under load84× faster
before
38 s
with PreMize
0.45 s
Concurrent chats served well8× capacity
before
2
with PreMize
16
Aggregate throughput6.3× throughput
before
20 tok/s
with PreMize
125 tok/s

From the production load-test campaign — not a projection.

11SOVEREIGNTY

Your documents never leave the machine.

100% on-premises

Models, index, and documents live on your hardware.

Air-gap capable

Once installed, it runs with no internet at all.

PDPA-aligned

Data residency by construction, with a full audit trail.

DGX Spark held in two hands — the whole platform at desk-side scale
95.1%
pass — 204-case live sweep
162
unit tests passing
0
open defects
PREMIZE — EFFICIENT LOCAL LLM

See it answer — with citations.

Contact sales

Appliance pricing, demos, and pilot programs for your organization.

[email protected]