SiamGPT Token Factory

Production AI inference for Thailand

One key.
Every model.

Frontier, open-source, and Thai-native models through one OpenAI-compatible API. We operate the GPUs, routing, and metering in Thailand. Your team integrates one endpoint and ships.

OpenAI, Anthropic, and Gemini-compatible · Billed per token · Tokens are AI inference output, not cryptocurrency.

DIFF app.py ONE LINE CHANGED
1client = OpenAI( 2- base_url="https://api.openai.com/v1", 2+ base_url="https://api.siamgpt.com/v1", 3 api_key="sgt-•••a1", 4) 5 6client.chat.completions.create( 7 model="siamllm-th", # or claude-sonnet-5, gpt-5.5 … 8 messages=[{"role": "user", "content": "สรุปสัญญาฉบับนี้"}], 9)
same SDK and payloads, key scoped to project support-bot
routed across providers with automatic failover
every request metered, attributed, and logged
One endpoint for
  • SiamLLM
  • OpenAI
  • Anthropic
  • Google
  • DeepSeek
  • Qwen
  • Meta Llama
  • Moonshot
  • Z.ai
  • Mistral
1API key per project, revocable in one click
3API styles: OpenAI, Anthropic, and Gemini
5Modalities: text, image, video, speech, embeddings
100%Of requests metered and attributed to a project

Managed inference

The token factory for commercial AI

SiamGPT Token Factory hosts and operates large language models for businesses through managed APIs and routing. You receive model outputs. We run the capacity, the serving stack, and the operations underneath.

01

Managed model APIs

Stable, versioned endpoints for commercial inference, compatible with the SDKs your team already uses.

02

Production token serving

We run the serving layer, capacity planning, observability, and reliability needed to produce tokens under real traffic.

03

Intelligent routing

Requests are weighed by health, latency, capacity, and cost, then routed with failover that switches in seconds.

04

Served in Thailand

Inference runs on NVIDIA GPUs operated in Thailand, so Thai data can stay in the country when your policy requires it.

Platform

Every way to buy inference, on one platform

Start pay-as-you-go, commit capacity once traffic is predictable, or isolate everything on hardware reserved for you. The API, keys, and bill stay the same.

Pay per token

Serverless inference

Call any catalog model on demand. No infrastructure to manage and no commitment.

Best for prototypes, new features, spiky traffic

Asynchronous

Batch inference

Submit large offline jobs and collect results later, at a lower rate than real-time calls.

Best for document backfills, evaluations, embeddings

Reserved + SLA

Provisioned throughput

Committed token capacity with a service level agreement, for production traffic that cannot queue.

Best for customer-facing assistants at steady volume

Single tenant

Dedicated endpoints

A model deployed on GPUs allocated to you alone, with predictable latency and your own configuration.

Best for latency-sensitive and custom models

On-premises

Private deployment

The full serving stack inside your data center or an isolated enclave, operated by our team.

Best for government, finance, regulated data

Model shaping

Fine-tuning & evaluation

Adapt SiamLLM or open models on your data, and measure quality before anything reaches production.

Best for domain language, house style, accuracy targets

Explore the platform

Governance

Who can call which model is decided by structure

Organizations, workspaces, and projects scope every key and every model. Permissions follow the structure, not individual people.

  • Three-level hierarchyOrganization, workspace, and project each carry their own members, roles, and credit budgets.
  • Model allowlistsA project can only call the model IDs enabled for it, checked before traffic reaches any provider.
  • Scoped keysEach key belongs to one project. Revoke it and that project's access ends everywhere at once.
Access structureClick a level

Project: holds the API keys and the model allowlist. This is where access is actually enforced.

Model allowlist3 / 5 enabled

Illustration. The allowlist applies before traffic reaches a provider.

Request log All2xxErrors
09:41:22siamllm-th20012.4k tok3.10 cr
project support-bot · key sgt-•••a1 · workspace Customer Experience
09:41:20gpt-5.52003.1k tok0.68 cr
09:41:18kling-v3-pro20042.00 cr
09:41:15gemini-3.5-flash4290.00 cr
09:41:11nemo-retriever-embed2000.8k tok0.04 cr

Illustrative data. Failed requests are logged and not charged.

Metering & billing

Every token has an owner

Usage from every provider is logged per request and rolled into one bill. Credits flow down the hierarchy, so budgets are set before the spend, not reconciled after it.

  • Per-request attributionModel, key, tokens, and credits on every call, traceable to a project and a department.
  • Credit allocationCredits move from the organization to workspaces and projects, with hard limits where you need them.
  • One invoiceSpend across every provider arrives as one statement, so finance deals with one counterparty.

A Monday morning

Developers, finance, and admins each see their own side

The same governance layer, opened by three people for three different reasons.

09:12 · DEVELOPER

Swap one base URL and a new model is live

Keep the OpenAI SDK. Scope comes from the project key, so there is no new authentication flow to build.

DIFFconfig.ts
1- model: "gpt-5.5", 1+ model: "siamllm-th", 2 baseURL: "https://api.siamgpt.com/v1",

Onboarding

From workload to tokens

Self-serve keys get you started in minutes. For committed and private capacity, a project runs in four steps.

Scope

We map the workload, target models, latency needs, throughput, and data residency.

Provision

We allocate and configure serving capacity for the agreed workload and delivery mode.

Serve

Your team integrates one API and receives model outputs, with keys and quotas per project.

Operate

We monitor routing, capacity, reliability, and token delivery as your demand changes.

Built in-house

SiamLLM, our Thai-native model

A large language model platform designed specifically for the Thai language. It sits in the catalog next to every other model, one model ID away.

  • Native understanding of Thai linguistic structures and nuances
  • Aware of Thai business etiquette, regulations, and social norms
  • Served on NVIDIA GPUs in Thailand, and powering the SiamGPT assistant

model: "siamllm-th" Meet SiamLLM

Solutions

Built for your operating reality

Stability, cost, compliance, and delivery requirements differ by team. Token Factory adapts from standard API access to fully private deployment.

01 · INTERNET & SAAS

AI features that survive peak traffic

One model API integrated once, with failover and per-project quotas and billing.

02 · AI STARTUPS

From prototype to launch without a platform team

Switch models freely on one protocol and pay only for the tokens you use.

03 · GOVERNMENT & FINANCE

Compliance and audit as defaults

Private deployment with isolated compute in Thailand and full-chain audit trails.

04 · RESEARCH & EDUCATION

Compare many models, trace every run

Unified multi-model calls, elastic concurrency, and complete call logs.

05 · CHANNEL PARTNERS

Your own AI service under your brand

Your domain and tenant, separate keys, and revenue-split billing.

06 · CONTENT & MEDIA

Text, image, and video on one bill

Multimodal models behind one interface, with high-concurrency batch generation.

Start building on Token Factory

Tell us your workload, target models, and expected traffic. We will scope capacity and get your team an API key.