Token FactoryPlatform

Platform

From GPU to token, operated for you

Inference, model shaping, governance, and billing on one platform. Pick the delivery mode that fits the workload today, and move between them without changing your integration.

Inference

Five ways to run a model

All five share the same API, keys, allowlists, and bill. What changes is how capacity is allocated underneath.

Pay per token

Serverless inference

The fastest way to run catalog models on demand. No infrastructure to manage and no long-term commitment. Capacity scales with your traffic automatically.

Best for prototypes, new features, spiky traffic

Asynchronous

Batch inference

Process large workloads asynchronously at a lower rate. Submit a file of requests, and collect results when the job finishes.

Best for document backfills, classification, embeddings, evaluations

Reserved + SLA

Provisioned throughput

Committed inference capacity with token-based pricing, reserved throughput, and a service level agreement. Drop-in compatible with serverless calls.

Best for production assistants at predictable volume

Single tenant

Dedicated endpoints

Deploy a model on GPUs allocated to you alone. Control batch size, context length, and scaling, with latency that does not depend on other tenants.

Best for latency-sensitive products, fine-tuned models

On-premises

Private deployment

Local deployment and private model hosting inside your facility or an isolated enclave, with dedicated operations from our team.

Best for government, finance, healthcare, sovereign data

Compute

GPU capacity

Training, inference, and batch capacity on NVIDIA GPUs for teams that need raw compute alongside managed endpoints.

Best for custom training runs and research workloads

Model shaping

Make a model yours, then prove it works

Fine-tune SiamLLM or open-weight models on your own data. Improve accuracy, reduce hallucinations, and control behavior, without managing training infrastructure.

Fine-tuning

Supervised and preference tuning on your corpus. The resulting model deploys to a dedicated endpoint behind the same API.

Evaluations

Measure model quality against your own task set, compare candidates side by side, and set a bar before production.

Optimization

Quantization, caching, and acceleration tuned to your scenarios, improving quality while cutting long-run cost.

Delivery stack

Seven capabilities in three layers

Resources keep it stable, the platform keeps it governed, and a single API keeps it simple.

Access

API

High-availability inference API

One entry point for chat, generation, retrieval, function calling, and embeddings, in OpenAI, Anthropic, and Gemini request styles. Streaming by default.

Platform

Model hub

Unified model catalog

Language, embedding, image, video, and speech models behind one protocol. Switch with a model ID, not a migration.

Control

Organizations & projects

Isolated workspaces with their own keys, roles, allowlists, quotas, logs, and bills.

Billing

Token metering

Usage-based pricing with tiered discounts and project-level split billing. Every call traceable to a cost.

Optimize

Model operations

Fine-tuning, quantization, caching, and acceleration for the scenarios that matter to you.

Foundation

Private

Private deployment

Local hosting for data security, compliance, and dedicated performance, with dedicated operations.

Resource

NVIDIA GPU pools

Serving, training, and batch capacity in Thailand, with elastic expansion as demand grows.

Delivery modes

Start small, commit when it makes sense

Pricing depends on model, volume, and delivery mode. Our team will quote against your actual workload.

Self-serve

Pay as you go

Top up credits and call any enabled model. Pay only for tokens used.

  • Serverless and batch inference
  • Full model catalog
  • Organizations, projects, scoped keys
  • Per-request logs and usage dashboard
Request access
Enterprise

Dedicated & private

Single-tenant endpoints or the full stack inside your own environment.

  • Dedicated GPUs or on-premises deployment
  • Data residency in Thailand
  • Fine-tuning and evaluation programs
  • Brand partnership under your own domain
Contact us

Architecture

The path a request takes

From application integration down to model and compute resources, every layer carries its own security, scheduling, governance, and failover.

Application

Your product

AI customer serviceKnowledge basesContent generationAgentsData analysis
Access

Unified access

OpenAI-compatible APIAnthropic-compatible APISDKsStreaming responses
Governance

Security & governance

AuthenticationTenant isolationModel allowlistsQuotas & rate limitsAudit & billing
Routing

Intelligent scheduling

Latency-aware routingLoad awarenessCircuit breakingAutomatic failover
Resource

Models & compute

SiamLLMPartner providersDedicated endpointsPrivate nodesNVIDIA GPU pools

Security matrix

Transport encryptionAccess whitelistsTenant data isolationFull-chain loggingMulti-node disaster recovery

Service level

More than a connection. A service level.

01

Multi-dimensional routing

Load, latency, capacity, health, and cost weighed together, with failover that switches in seconds.

02

Enterprise governance

Tenants, permissions, keys, quotas, audit trails, billing, and risk control, end to end.

03

Full stack, GPU to token

Infrastructure and model operations optimized as one path, so customization and response are fast.

04

One protocol, every vendor

Standardized calls across providers remove duplicate development and migration cost.

05

Three delivery modes

Public cloud, brand partnership under your own domain, or fully private deployment.

06

Cost-aware scheduling

Pay for what you use, while scheduling keeps reducing long-run spend.

Infrastructure

NVIDIA Technology Implementation

Token Factory runs on the same infrastructure plan as SiamLLM, so the capacity you call through the API is capacity we operate ourselves.

Computing Infrastructure

We plan to deploy NVIDIA DGX B200 for training and serving large AI models. Our products run on NVIDIA GPUs.

AI Development Frameworks

  • For data ingestion, we implement the NVIDIA NeMo Retriever, followed by a data curation process. We then utilize the synthetic data generation pipelines from NVIDIA NeMo Curator.
  • The Query Orchestrator manages initial processing, creates pre-filters, and builds context through vector search. It then leverages NVIDIA NeMo Retriever and NIM microservices – including Deepseek R1, Riva NMT, Llama 3.1, NeMo Retriever reranking, and NeMo Retriever embedding NIM – to produce accurate and contextually relevant responses.

Tell us about your workload

Share your target models, expected traffic, and data residency needs. We will recommend a delivery mode and quote against it.