Serverless inference
The fastest way to run catalog models on demand. No infrastructure to manage and no long-term commitment. Capacity scales with your traffic automatically.
Best for prototypes, new features, spiky traffic
Token Factory › Platform
Platform
Inference, model shaping, governance, and billing on one platform. Pick the delivery mode that fits the workload today, and move between them without changing your integration.
Inference
All five share the same API, keys, allowlists, and bill. What changes is how capacity is allocated underneath.
The fastest way to run catalog models on demand. No infrastructure to manage and no long-term commitment. Capacity scales with your traffic automatically.
Best for prototypes, new features, spiky traffic
Process large workloads asynchronously at a lower rate. Submit a file of requests, and collect results when the job finishes.
Best for document backfills, classification, embeddings, evaluations
Committed inference capacity with token-based pricing, reserved throughput, and a service level agreement. Drop-in compatible with serverless calls.
Best for production assistants at predictable volume
Deploy a model on GPUs allocated to you alone. Control batch size, context length, and scaling, with latency that does not depend on other tenants.
Best for latency-sensitive products, fine-tuned models
Local deployment and private model hosting inside your facility or an isolated enclave, with dedicated operations from our team.
Best for government, finance, healthcare, sovereign data
Training, inference, and batch capacity on NVIDIA GPUs for teams that need raw compute alongside managed endpoints.
Best for custom training runs and research workloads
Model shaping
Fine-tune SiamLLM or open-weight models on your own data. Improve accuracy, reduce hallucinations, and control behavior, without managing training infrastructure.
Supervised and preference tuning on your corpus. The resulting model deploys to a dedicated endpoint behind the same API.
Measure model quality against your own task set, compare candidates side by side, and set a bar before production.
Quantization, caching, and acceleration tuned to your scenarios, improving quality while cutting long-run cost.
Delivery stack
Resources keep it stable, the platform keeps it governed, and a single API keeps it simple.
Access
One entry point for chat, generation, retrieval, function calling, and embeddings, in OpenAI, Anthropic, and Gemini request styles. Streaming by default.
Platform
Language, embedding, image, video, and speech models behind one protocol. Switch with a model ID, not a migration.
Isolated workspaces with their own keys, roles, allowlists, quotas, logs, and bills.
Usage-based pricing with tiered discounts and project-level split billing. Every call traceable to a cost.
Fine-tuning, quantization, caching, and acceleration for the scenarios that matter to you.
Foundation
Local hosting for data security, compliance, and dedicated performance, with dedicated operations.
Serving, training, and batch capacity in Thailand, with elastic expansion as demand grows.
Delivery modes
Pricing depends on model, volume, and delivery mode. Our team will quote against your actual workload.
Top up credits and call any enabled model. Pay only for tokens used.
Reserved token capacity with a service level agreement for steady traffic.
Single-tenant endpoints or the full stack inside your own environment.
Architecture
From application integration down to model and compute resources, every layer carries its own security, scheduling, governance, and failover.
Security matrix
Service level
Load, latency, capacity, health, and cost weighed together, with failover that switches in seconds.
Tenants, permissions, keys, quotas, audit trails, billing, and risk control, end to end.
Infrastructure and model operations optimized as one path, so customization and response are fast.
Standardized calls across providers remove duplicate development and migration cost.
Public cloud, brand partnership under your own domain, or fully private deployment.
Pay for what you use, while scheduling keeps reducing long-run spend.
Infrastructure
Token Factory runs on the same infrastructure plan as SiamLLM, so the capacity you call through the API is capacity we operate ourselves.
We plan to deploy NVIDIA DGX B200 for training and serving large AI models. Our products run on NVIDIA GPUs.
Share your target models, expected traffic, and data residency needs. We will recommend a delivery mode and quote against it.