How I would run pre-sales PoCs and escalations for APAC accounts, built on live Runpod data, plus a working workbench that sizes a model against live GPU prices, maps APAC capacity, debugs a Serverless endpoint, triages logs, checks REST v1 code and generates a PoC kit.
Pre-sales engineer for high-spend APAC accounts and the escalation engineer for APAC hours, paired with the Account Executive APAC. Wins are measured in onboarded spend and resolved escalations.
APAC demand outruns APAC capacity: three listed Asia-Pacific datacenters, thin H100 and H200 stock, and no S3 API endpoint in the region. Placement advice is half of every APAC deal.
A six-tab workbench on Runpod's live API: sizing and pricing, APAC capacity, an endpoint doctor, log triage, a REST v1 to v2 checker and a generated PoC kit.
| Takeaway | What it means for an APAC FDE |
|---|---|
| Placement decides the PoC | A network volume pins every worker to one datacenter. In a low-stock Asian datacenter with one GPU type, a PoC endpoint can sit at zero workers on demo day. |
| 15 November is a hard date | REST API v1 retires on 2026-11-15 and GraphQL follows in early 2027. Every APAC account with provisioning scripts needs a check in the first weeks. |
| Serverless versus Pod is a sizing answer | On H100, a Serverless flex worker ($4.79/hr) beats an always-on Secure Pod ($3.49/hr) until it is busy about 73% of the day. The PoC should measure that number. |
| Escalations repeat | Most tickets map to a short list of signatures: OOM, KV cache, gated models, image architecture, CUDA mismatch, throttled workers. Each has a known first check and owner. |
A developer cloud that grew from self-serve GPUs into a full platform. The revenue team now sells to enterprises and frontier-model teams in production.
| Product | What it is | Where it shows up in FDE work |
|---|---|---|
| Pods | GPU containers by the hour, Secure Cloud or Community Cloud | Training, fine-tuning, steady inference; the place to reproduce a Serverless bug |
| Serverless (queue) | Per-second workers behind /run, /runsync, /status; flex and active workers | Most inference PoCs; cold starts, throttling and queue depth are the usual tickets |
| Serverless (load balancer) | Direct HTTP to workers on custom paths | Customers who bring their own server (vLLM, SGLang, custom APIs) |
| vLLM worker | Runpod's worker image with an OpenAI-compatible route | The default LLM PoC; configured through environment variables |
| Instant and Reserved Clusters | Multi-node GPU clusters, up to 64 GPUs on demand, reserved by contract | Training deals and large commits, quoted with the AE |
| Public Endpoints | Hosted models priced per token, image or second | Fast first value for customers who want an API before a deployment |
| Network volumes and S3 API | Persistent storage per datacenter; S3-compatible access in 15 datacenters | Data placement, cold-start reduction, and the main cause of region lock |
| Flash, Hub, runpodctl | Python SDK for endpoints from local code, a repo catalogue, and the CLI | Developer experience; an open runpodctl issue (#345) asks when its pod commands move off REST v1 |
Read from Runpod's public datacenter API on 5 October 2026. The demo's APAC tab refreshes this live and ranks datacenters by distance from the customer's city.
| Datacenter | Status | GPU stock | Planning note |
|---|---|---|---|
| AP-JP-1 (Japan) | listed | H200 SXM low, H100 SXM low | The only APAC site with Hopper H200 and the API's storage flag. First choice for Japan and Korea |
| AP-IN-1 (India) | listed | H100 SXM low | Nearest H100 to Southeast Asia by distance; AP-IN-2 also shows H100 but is unlisted |
| OC-AU-1 (Australia) | listed | L40S low | 48 GB inference tier for Australia and New Zealand |
| SEA-SG-1 (Singapore) | unlisted | none shown | Present in the API with no GPU stock; worth asking about plans |
| Malaysia, Korea | none | n/a | Customers here run in Japan, India, or outside APAC; local sovereign clouds compete on residency |
List several pools per Serverless endpoint and leave datacenter pinning off unless residency requires it.
None of the 15 S3 API datacenters is in APAC. Bulk uploads for an Asian workload go through a Pod, runpodctl, or an external bucket.
Kuala Lumpur to AP-IN-1 is about 3,600 km if the site is in Mumbai (RTT floor about 54 ms); to AP-JP-1 about 5,300 km if in Tokyo (about 80 ms). Fine for streamed chat, costly for chatty sequential calls.
H100 on-demand from each vendor's public pricing page, October 2026. APAC-native clouds win on residency and local billing; Runpod wins on per-second Serverless and self-serve breadth.
| Provider | Positioning | H100 on-demand | Gap versus Runpod |
|---|---|---|---|
| Runpod | Developer cloud: Pods, Serverless, Clusters, Public Endpoints | SXM $3.49/hr Secure, $2.69 Community; Serverless $4.79/hr | Baseline |
| CoreWeave | Hyperscale AI cloud for labs and large enterprise | 8x HGX $49.24/hr (about $6.16/GPU) | Contract-led 8-GPU nodes; Runpod wins single-GPU self-serve and per-second billing |
| Lambda | Training clouds and 1-Click Clusters | 1x SXM $4.29/hr | No serverless inference product |
| Modal | Python-first serverless GPU | About $3.95/hr, billed per second | Closest developer-experience rival to Serverless and Flash; no persistent Pods |
| Together AI | Hosted model APIs, fine-tuning, clusters | Instant cluster $3.99/hr | Leads on hosted APIs; Runpod leads on bring-your-own-container |
| Vast.ai | Peer GPU marketplace | From about $1.79/hr | Lower floor, weaker reliability and compliance story |
| Fireworks | Optimised hosted inference | $8.00/hr on-demand deployment | Owns serving optimisations; Runpod is cheaper raw compute |
| Baseten | Managed enterprise model serving | About $6.50/hr | Premium managed tier |
| Nebius | Full-stack AI cloud | $3.85/hr, rising to $4.50 from 1 Oct 2026 | Strong reserved clusters; no APAC-first footprint |
| YTL AI Cloud | Malaysian sovereign AI cloud, Johor | GB200 RM 68.80/hr listed | Malaysian data residency; Runpod has no Malaysian region |
| Singtel RE:AI | Singapore telco sovereign AI cloud | Not public | Enterprise and public-sector channel in Singapore |
| Sakura Internet | Japanese government-backed GPU cloud | Not verified | Yen billing and Japanese residency; Runpod's answer is AP-JP-1 |
| KT Cloud, Naver Cloud, Elice | Korean sovereign and education clouds | Not verified | Korean-language support and procurement channels |
Found while building the demo, each checked twice against the live pages on 5 October 2026. Small items, and the kind customers quote back in a sales call.
| Finding | Evidence | Suggested fix |
|---|---|---|
| Granite 4.0 H Small priced 10x apart | Pricing page, Public Endpoints: "$1.00 per 1m tokens". Docs model list and model page: "$10.00 per 1M tokens" | Pick one source of truth and generate the other from it |
| A retired-looking model at a placeholder price | Pricing page lists "Deep Cogito v2 Llama 70B" at "$0.00001 per 1m tokens"; the docs model list does not include it | Remove or reprice the row |
| S3 docs use a host and region that do not exist | --endpoint-url https://storage.datacenter.runpod.io in the timeout example, and region ca-qc-1 in the boto3 help text; neither is in the S3 datacenter table | Use the documented s3api-<dc>.runpod.io pattern and a real datacenter ID |
| Global volumes page links to a 404 | The link to /storage/globalvolume-pods returns 404; the page lives at /storage/globalvolume/globalvolume-pods. The same page says any Pod or worker can mount a global volume, then that Serverless is unsupported | Fix the link and state Pods only once |
| B300 memory listed two ways | Same pricing page: "288 GB" in the Pods table, "280 GB" in the Serverless table; API and docs say 288 | Align on 288 |
| "Llama 3 7B" | Serverless 48 GB tier copy: "Extreme inference throughput on LLMs like Llama 3 7B". Llama 3 shipped 8B and 70B | Change to Llama 3.1 8B |
Building note: the GPU catalog API returns a price of 0.5 for clouds a GPU is not offered on (for example a community price on GPUs with communityCloud: false). The demo reads the cloud flag before showing a price; any customer script that skips the flag will quote wrong numbers.
| JD duty | How I would do it | Detailed in |
|---|---|---|
| Sales meetings, architecture, PoCs for high-spend prospects | Discovery call with a sizing sheet; architecture written down with GPU, datacenter and Serverless or Pod choice; PoC with success criteria agreed on day 0 and a load test the customer can rerun | §05, demo tabs 1, 2, 6 |
| Troubleshoot critical escalations | Health counters first, then logs, then reproduce on a Pod with the same image. Reply with cause, fix and owner in one message | §06, demo tabs 3, 4 |
| Code analysis, scripts, log analysis, remote access | Read the customer's handler and Dockerfile; scripted repro against the job API; signature-based log triage before deep debugging | Demo tabs 4, 5 |
| Communicate across engineering, sales, supply, product | One escalation format: symptom, evidence, customer impact, ask, deadline. Supply gets capacity asks with datacenter and GPU named | §06 |
| Assist support, work with infra on GPU servers | Take APAC-hours escalations from L2; hand infra a host ID, timestamp and failing line, never a summary | §06 |
| Relay feedback into product | Monthly ranked list from APAC tickets and PoCs, each with account names and spend at stake | §08 |
| Docs, KB, webinars, demos | Write the fix once as a KB article after the second identical ticket; an APAC placement guide; a REST v2 migration webinar for APAC hours | §07, §08 |
Each stage has one output and one exit test, so the AE always knows where a deal stands.
| Stage | What I do | Output | Exit test |
|---|---|---|---|
| Discovery | Model, context length, concurrency, traffic shape, latency target, residency, current spend and provider | Sizing sheet | Numbers for every input, signed off by their engineer |
| Architecture | GPU and count, datacenter shortlist, Serverless or Pod or cluster, storage plan | One-page design with costs | Customer agrees the design and the PoC budget |
| PoC | Deploy from the v2 API, run their load test, measure TTFT, throughput, cold start and cost | Results against the day-0 criteria | All criteria pass, or a written gap with a fix date |
| Production | Active workers or reserved capacity sized from PoC data; alerts on queue delay and unhealthy workers | Runbook and handover to support | Two weeks of production traffic without an escalation |
| Expansion | Review spend and utilisation monthly; next workload (fine-tuning, a second model, batch jobs) | Expansion plan with the AE | Second workload in PoC |
Five escalations I would expect in the first month of APAC hours, built from Runpod's docs, GitHub issues and status history.
| Symptom | First checks | Likely cause and fix | Escalate to |
|---|---|---|---|
| PoC endpoint in Japan stuck at 0 workers before a demo | Live stock for AP-JP-1; GPU types on the endpoint; attached network volume | Single volume pins workers to one low-stock datacenter. Widen GPU pools, hold one active worker for the demo, add a second volume elsewhere | Supply for an AP-JP-1 reservation on a real deal; AE for reserved pricing |
| vLLM endpoint returns 404 model not found after an upgrade | /openai/v1/models versus the client's model; worker image tag | Served-name mismatch (worker-vllm #310 was one such regression, closed June 2026). Set OPENAI_SERVED_MODEL_NAME_OVERRIDE or pin the tag | worker-vllm owner if a new release regresses |
| Jobs sit IN_QUEUE while workers idle, or the endpoint scaled itself down | /health counters; SDK version in the image; worker logs | runpod-python 1.7.11 to 1.10.0 could stop workers pulling jobs on volume endpoints: upgrade to 1.10.1+. Repeated unhealthy workers trigger scale-down: fix the crash first | Engineering only if the upgrade does not fix it |
| Provisioning script will break on 15 November | Grep for rest.runpod.io/v1, flat bodies, bare-array parsing, /billing/endpoints | REST v1 retirement. Job API calls on api.runpod.ai/v2 are unaffected; the billing path is the easy one to miss | Product if they rely on Pod reset, which has no v2 action |
| Billed twice for long jobs; tax on credit purchases | Job IDs, execution time against executionTimeout, retries; billing country | Long jobs retried after timeout re-run without idempotency. Set timeout and TTL to the real job; dedupe in the handler. Tax depends on billing country; a valid tax ID may exempt a business | Billing team for a credit review, with job IDs attached |
The demo's endpoint doctor and log triage tabs implement the first checks for these and eight more log signatures.
The most time-boxed work an APAC FDE inherits. From Runpod's migration guide and the v2 OpenAPI schema.
| Change | v1 to v2 | What breaks if missed |
|---|---|---|
| Base URL | rest.runpod.io/v1 to api.runpod.io/v2 | Every call, on the retirement date |
| Serverless billing path | /billing/endpoints to /v2/billing/serverless | /v2/billing/endpoints exists and returns Public Endpoint billing, so a naive rename returns the wrong data without an error |
| Pod lifecycle | /pods/{id}/stop to POST /v2/pods/{id}/action with {"action":"stop"} | reset has no v2 action; nightly reset jobs need a decision |
| Create bodies | Flat fields to nested gpu, workers, scaling, mounts; endpoint type required | 400 errors on create |
| List responses | Bare arrays to {"pods":[...]} style wrappers | Loops over the response iterate keys or fail |
| Errors | {"message"} to RFC 9457 title, status, detail | Error handlers raise their own exceptions |
| GPU selection on PATCH | gpu.pools and excludedTypes are replaced together | Sending pools alone clears exclusions and can land workers on unwanted GPUs |
The v2 catalog endpoints need an API key, while today's GraphQL catalog query does not. Price trackers and scripts that read prices without a key will need one after GraphQL retires.
List APAC accounts with v1 traffic, run each integration through a checker, and book fixes before 15 November. Protects spend at no acquisition cost.
Every APAC PoC ships with multiple GPU pools, no datacenter pin unless required, and one active worker on demo days. Removes the most common reason a PoC fails in front of a CTO.
One page on where to run from each APAC city, how data gets there without an APAC S3 endpoint, and when residency rules force a choice. Reused in every deal.
For accounts above a spend line, compare busy share with the break-even monthly. Moving the right workloads to active workers or reserved capacity lowers their bill and raises commitment.
Second identical ticket becomes a KB article; fifth becomes a docs change or a product request with spend attached.
A monthly note to supply: APAC demand by GPU type and datacenter, lost or delayed deals, and the residency asks behind them.
| Window | Do | Measured by |
|---|---|---|
| Days 1 to 30 | Shadow L2 and the AE; run every product end to end on my own account; v1 sweep for APAC accounts; take APAC-hours escalations | Every APAC v1 integration identified, owners booked before 15 November |
| Days 31 to 60 | Own PoCs for the AE's pipeline; ship the PoC template and placement guide; first KB articles from repeat tickets | PoCs run against day-0 criteria; time to first PoC result |
| Days 61 to 90 | First production handovers; monthly feedback list to product; first APAC webinar or partner workshop | Onboarded spend from PoCs; escalation time to resolution in APAC hours |
Public sources only, read on 5 October 2026 (public job posting, 2026). Live figures come from Runpod's public GraphQL API. The demo is my own prototype built for this application.
Company: Series A post · $120M ARR note · TechCrunch
Pricing: runpod.io/pricing · Serverless pricing · Public Endpoints models
Serverless: operations · worker states · troubleshooting
vLLM: environment variables · worker-vllm #310 · runpod-python #402
API: Migrate from API v1 · v2 OpenAPI schema · release notes
Storage: S3-compatible API · global volumes
Status: uptime.runpod.io
Competitors: CoreWeave · Lambda · Modal · Together · Nebius · YTL AI
Model configs: Hugging Face config.json for each preset in the demo
Independent homework for the Runpod Forward Deployed Engineer APAC role · 2026 · edwardtay.com