M Mlandi AI
Docs Pricing Status About Console
Developers

API Documentation

Base URL: https://api.mlandi.com/v1 · OpenAI-compatible

Call China's leading models through a single OpenAI-compatible endpoint. Change your base URL and key and your existing SDK code keeps working — no Chinese phone number, no CNY billing, no separate consoles.

Quick start

If your code already talks to the OpenAI SDK, it works here. Change two lines: the base URL and the key.

Endpoint: https://api.mlandi.com/v1  ·  Auth: Authorization: Bearer sk-...

cURL
curl https://api.mlandi.com/v1/chat/completions \
  -H "Authorization: Bearer $MLANDI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-chat",
    "messages": [{"role": "user", "content": "Explain TCP handshake in one paragraph."}]
  }'
Python (openai SDK)
from openai import OpenAI

client = OpenAI(
    base_url="https://api.mlandi.com/v1",
    api_key="sk-...",
)

resp = client.chat.completions.create(
    model="deepseek-chat",
    messages=[{"role": "user", "content": "Hello from Manila"}],
)
print(resp.choices[0].message.content)
Node.js
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.mlandi.com/v1",
  apiKey: process.env.MLANDI_API_KEY,
});

const resp = await client.chat.completions.create({
  model: "deepseek-chat",
  messages: [{ role: "user", content: "Hello from Dubai" }],
});

console.log(resp.choices[0].message.content);
Streaming
stream = client.chat.completions.create(
    model="deepseek-chat",
    messages=[{"role": "user", "content": "Write a haiku about latency"}],
    stream=True,
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Authentication

  • Create a key in the console and send it as a bearer token.
  • Never expose keys in browser or mobile code — call us from your backend.
  • Use a separate key per environment so you can rotate safely.
  • Set a per-key spend cap to bound the damage of a leak. Revocation takes effect within seconds.

Models

All models are served through each provider's official API. We do not run reverse-engineered endpoints, shared consumer accounts, or unofficial proxies.

Model IDProviderContextBest for
deepseek-chatDeepSeek128KGeneral chat and coding, lowest cost per token
deepseek-reasonerDeepSeek128KMulti-step reasoning, math, agent planning
qwen-maxAlibaba Qwen128KStrong multilingual, enterprise workloads
qwen-plusAlibaba Qwen128KBalanced quality and cost
glm-4.6Zhipu GLM200KLong context, agentic tool use
kimi-k2Moonshot AI128KLong documents, agent workflows
doubao-proByteDance256KHigh throughput, Chinese-heavy workloads
hunyuan-turboTencent128KFast responses, cost-sensitive apps

Model availability follows each provider's roadmap. Query GET /v1/models for the live list bound to your key.

curl https://api.mlandi.com/v1/models -H "Authorization: Bearer $MLANDI_API_KEY"

Endpoints

MethodPathDescription
POST/v1/chat/completionsChat and instruction following, streaming supported
GET/v1/modelsModels available to your key
POST/v1/embeddingsEmbeddings where the upstream provider supports them

Function calling and JSON mode follow the OpenAI schema and pass through to the provider. Support varies per model.

Errors

StatusMeaningWhat to do
400Malformed request or unsupported parameterCheck the body against the OpenAI schema
401Missing or invalid keyVerify the Authorization header
402Insufficient balanceTop up; requests resume automatically
404Unknown model or routeCall GET /v1/models
429Rate limit or upstream throttleBack off exponentially; honour Retry-After
502 / 504Upstream error or timeoutRetry; check the status page

Retry tip: for 429 and 5xx, back off exponentially with jitter. Set a client timeout of at least 120 seconds — reasoning models legitimately take that long.

Regions and latency

Asia Pacific · default

api.mlandi.com/v1

Served from Bangkok. Lowest latency for South East Asia, South Asia, East Asia and the Middle East.

Americas · edge

us-west.mlandi.com/v1

US West forwarding edge for North and South America, removing one long round trip per call.

Latency is measured continuously and published on the status page as p50 and p95, not marketing averages.

Works with your existing tools

Claude Code

Point ANTHROPIC_BASE_URL at us to route coding sessions through your own balance.

Cursor / Continue

Set the custom OpenAI base URL and paste your key.

Open WebUI

Add an OpenAI-compatible connection, then pick models from the dropdown.

LangChain

Pass base_url to the OpenAI-compatible wrapper. No adapter needed.

n8n / Dify

Use the OpenAI credential type with a custom endpoint.

Your product

Drop-in SDK swap. Ship to production in minutes.

Limits and fair use

  • Default 60 requests per minute per key; higher tiers on request.
  • Maximum request body 512 KB. Chunk larger workloads or use a long-context model.
  • Streaming responses are not buffered, keeping time-to-first-token low.
  • Abuse, spam generation and illegal content are prohibited under our Acceptable Use Policy.
Support: support@mlandi.com — we answer within one business day, sooner for anything affecting production traffic. Live incident updates are published on the status page rather than buried in a chat thread.
Mlandi AI · www.mlandi.com
API Docs · Pricing · Status · About · Acceptable Use Policy · Terms of Service · Privacy Policy
Privacy: privacy@mlandi.com · Support: support@mlandi.com