FDE HOMEWORK ·Forward Deployed Engineer APAC, Runpod
ⓘ Independent job-application page. Not affiliated with, endorsed by, or operated by Runpod. Public information as of 5 October 2026.

Homework for my application to Runpod Forward Deployed Engineer APAC.

How I would run pre-sales PoCs and escalations for APAC accounts, built on live Runpod data, plus a working workbench that sizes a model against live GPU prices, maps APAC capacity, debugs a Serverless endpoint, triages logs, checks REST v1 code and generates a PoC kit.

1M+
developers on the platform (company figure)
3 of 34
listed datacenters in Asia-Pacific (live API, 5 Oct)
41 days
to REST API v1 retirement on 15 Nov 2026
6
public pricing and docs inconsistencies found while building
00

Summary

The seat

Pre-sales engineer for high-spend APAC accounts and the escalation engineer for APAC hours, paired with the Account Executive APAC. Wins are measured in onboarded spend and resolved escalations.

The APAC constraint

APAC demand outruns APAC capacity: three listed Asia-Pacific datacenters, thin H100 and H200 stock, and no S3 API endpoint in the region. Placement advice is half of every APAC deal.

The work sample

A six-tab workbench on Runpod's live API: sizing and pricing, APAC capacity, an endpoint doctor, log triage, a REST v1 to v2 checker and a generated PoC kit.

TakeawayWhat it means for an APAC FDE
Placement decides the PoCA network volume pins every worker to one datacenter. In a low-stock Asian datacenter with one GPU type, a PoC endpoint can sit at zero workers on demo day.
15 November is a hard dateREST API v1 retires on 2026-11-15 and GraphQL follows in early 2027. Every APAC account with provisioning scripts needs a check in the first weeks.
Serverless versus Pod is a sizing answerOn H100, a Serverless flex worker ($4.79/hr) beats an always-on Secure Pod ($3.49/hr) until it is busy about 73% of the day. The PoC should measure that number.
Escalations repeatMost tickets map to a short list of signatures: OOM, KV cache, gated models, image architecture, CUDA mismatch, throttled workers. Each has a known first check and owner.
01

Runpod in 2026

A developer cloud that grew from self-serve GPUs into a full platform. The revenue team now sells to enterprises and frontier-model teams in production.

$100M
Series A, June 2026, led by Summit Partners
$120M
ARR run rate, January 2026 (company post)
20B+
inference requests processed (company figure)
52
datacenters in the API, 34 listed
ProductWhat it isWhere it shows up in FDE work
PodsGPU containers by the hour, Secure Cloud or Community CloudTraining, fine-tuning, steady inference; the place to reproduce a Serverless bug
Serverless (queue)Per-second workers behind /run, /runsync, /status; flex and active workersMost inference PoCs; cold starts, throttling and queue depth are the usual tickets
Serverless (load balancer)Direct HTTP to workers on custom pathsCustomers who bring their own server (vLLM, SGLang, custom APIs)
vLLM workerRunpod's worker image with an OpenAI-compatible routeThe default LLM PoC; configured through environment variables
Instant and Reserved ClustersMulti-node GPU clusters, up to 64 GPUs on demand, reserved by contractTraining deals and large commits, quoted with the AE
Public EndpointsHosted models priced per token, image or secondFast first value for customers who want an API before a deployment
Network volumes and S3 APIPersistent storage per datacenter; S3-compatible access in 15 datacentersData placement, cold-start reduction, and the main cause of region lock
Flash, Hub, runpodctlPython SDK for endpoints from local code, a repo catalogue, and the CLIDeveloper experience; an open runpodctl issue (#345) asks when its pod commands move off REST v1
02

APAC capacity, from live data

Read from Runpod's public datacenter API on 5 October 2026. The demo's APAC tab refreshes this live and ranks datacenters by distance from the customer's city.

DatacenterStatusGPU stockPlanning note
AP-JP-1 (Japan)listedH200 SXM low, H100 SXM lowThe only APAC site with Hopper H200 and the API's storage flag. First choice for Japan and Korea
AP-IN-1 (India)listedH100 SXM lowNearest H100 to Southeast Asia by distance; AP-IN-2 also shows H100 but is unlisted
OC-AU-1 (Australia)listedL40S low48 GB inference tier for Australia and New Zealand
SEA-SG-1 (Singapore)unlistednone shownPresent in the API with no GPU stock; worth asking about plans
Malaysia, Koreanonen/aCustomers here run in Japan, India, or outside APAC; local sovereign clouds compete on residency
Hedge GPU pools

List several pools per Serverless endpoint and leave datacenter pinning off unless residency requires it.

Plan data paths

None of the 15 S3 API datacenters is in APAC. Bulk uploads for an Asian workload go through a Pod, runpodctl, or an external bucket.

Distance is a floor

Kuala Lumpur to AP-IN-1 is about 3,600 km if the site is in Mumbai (RTT floor about 54 ms); to AP-JP-1 about 5,300 km if in Tokyo (about 80 ms). Fine for streamed chat, costly for chatty sequential calls.

03

Competitive map

H100 on-demand from each vendor's public pricing page, October 2026. APAC-native clouds win on residency and local billing; Runpod wins on per-second Serverless and self-serve breadth.

ProviderPositioningH100 on-demandGap versus Runpod
RunpodDeveloper cloud: Pods, Serverless, Clusters, Public EndpointsSXM $3.49/hr Secure, $2.69 Community; Serverless $4.79/hrBaseline
CoreWeaveHyperscale AI cloud for labs and large enterprise8x HGX $49.24/hr (about $6.16/GPU)Contract-led 8-GPU nodes; Runpod wins single-GPU self-serve and per-second billing
LambdaTraining clouds and 1-Click Clusters1x SXM $4.29/hrNo serverless inference product
ModalPython-first serverless GPUAbout $3.95/hr, billed per secondClosest developer-experience rival to Serverless and Flash; no persistent Pods
Together AIHosted model APIs, fine-tuning, clustersInstant cluster $3.99/hrLeads on hosted APIs; Runpod leads on bring-your-own-container
Vast.aiPeer GPU marketplaceFrom about $1.79/hrLower floor, weaker reliability and compliance story
FireworksOptimised hosted inference$8.00/hr on-demand deploymentOwns serving optimisations; Runpod is cheaper raw compute
BasetenManaged enterprise model servingAbout $6.50/hrPremium managed tier
NebiusFull-stack AI cloud$3.85/hr, rising to $4.50 from 1 Oct 2026Strong reserved clusters; no APAC-first footprint
YTL AI CloudMalaysian sovereign AI cloud, JohorGB200 RM 68.80/hr listedMalaysian data residency; Runpod has no Malaysian region
Singtel RE:AISingapore telco sovereign AI cloudNot publicEnterprise and public-sector channel in Singapore
Sakura InternetJapanese government-backed GPU cloudNot verifiedYen billing and Japanese residency; Runpod's answer is AP-JP-1
KT Cloud, Naver Cloud, EliceKorean sovereign and education cloudsNot verifiedKorean-language support and procurement channels
★

Public-facts findings

Found while building the demo, each checked twice against the live pages on 5 October 2026. Small items, and the kind customers quote back in a sales call.

FindingEvidenceSuggested fix
Granite 4.0 H Small priced 10x apartPricing page, Public Endpoints: "$1.00 per 1m tokens". Docs model list and model page: "$10.00 per 1M tokens"Pick one source of truth and generate the other from it
A retired-looking model at a placeholder pricePricing page lists "Deep Cogito v2 Llama 70B" at "$0.00001 per 1m tokens"; the docs model list does not include itRemove or reprice the row
S3 docs use a host and region that do not exist--endpoint-url https://storage.datacenter.runpod.io in the timeout example, and region ca-qc-1 in the boto3 help text; neither is in the S3 datacenter tableUse the documented s3api-<dc>.runpod.io pattern and a real datacenter ID
Global volumes page links to a 404The link to /storage/globalvolume-pods returns 404; the page lives at /storage/globalvolume/globalvolume-pods. The same page says any Pod or worker can mount a global volume, then that Serverless is unsupportedFix the link and state Pods only once
B300 memory listed two waysSame pricing page: "288 GB" in the Pods table, "280 GB" in the Serverless table; API and docs say 288Align on 288
"Llama 3 7B"Serverless 48 GB tier copy: "Extreme inference throughput on LLMs like Llama 3 7B". Llama 3 shipped 8B and 70BChange to Llama 3.1 8B

Building note: the GPU catalog API returns a price of 0.5 for clouds a GPU is not offered on (for example a community price on GPUs with communityCloud: false). The demo reads the cloud flag before showing a price; any customer script that skips the flag will quote wrong numbers.

04

JD duties, my plan

JD dutyHow I would do itDetailed in
Sales meetings, architecture, PoCs for high-spend prospectsDiscovery call with a sizing sheet; architecture written down with GPU, datacenter and Serverless or Pod choice; PoC with success criteria agreed on day 0 and a load test the customer can rerun§05, demo tabs 1, 2, 6
Troubleshoot critical escalationsHealth counters first, then logs, then reproduce on a Pod with the same image. Reply with cause, fix and owner in one message§06, demo tabs 3, 4
Code analysis, scripts, log analysis, remote accessRead the customer's handler and Dockerfile; scripted repro against the job API; signature-based log triage before deep debuggingDemo tabs 4, 5
Communicate across engineering, sales, supply, productOne escalation format: symptom, evidence, customer impact, ask, deadline. Supply gets capacity asks with datacenter and GPU named§06
Assist support, work with infra on GPU serversTake APAC-hours escalations from L2; hand infra a host ID, timestamp and failing line, never a summary§06
Relay feedback into productMonthly ranked list from APAC tickets and PoCs, each with account names and spend at stake§08
Docs, KB, webinars, demosWrite the fix once as a KB article after the second identical ticket; an APAC placement guide; a REST v2 migration webinar for APAC hours§07, §08
05

Discovery to production

Each stage has one output and one exit test, so the AE always knows where a deal stands.

StageWhat I doOutputExit test
DiscoveryModel, context length, concurrency, traffic shape, latency target, residency, current spend and providerSizing sheetNumbers for every input, signed off by their engineer
ArchitectureGPU and count, datacenter shortlist, Serverless or Pod or cluster, storage planOne-page design with costsCustomer agrees the design and the PoC budget
PoCDeploy from the v2 API, run their load test, measure TTFT, throughput, cold start and costResults against the day-0 criteriaAll criteria pass, or a written gap with a fix date
ProductionActive workers or reserved capacity sized from PoC data; alerts on queue delay and unhealthy workersRunbook and handover to supportTwo weeks of production traffic without an escalation
ExpansionReview spend and utilisation monthly; next workload (fine-tuning, a second model, batch jobs)Expansion plan with the AESecond workload in PoC
06

Escalation playbook

Five escalations I would expect in the first month of APAC hours, built from Runpod's docs, GitHub issues and status history.

SymptomFirst checksLikely cause and fixEscalate to
PoC endpoint in Japan stuck at 0 workers before a demoLive stock for AP-JP-1; GPU types on the endpoint; attached network volumeSingle volume pins workers to one low-stock datacenter. Widen GPU pools, hold one active worker for the demo, add a second volume elsewhereSupply for an AP-JP-1 reservation on a real deal; AE for reserved pricing
vLLM endpoint returns 404 model not found after an upgrade/openai/v1/models versus the client's model; worker image tagServed-name mismatch (worker-vllm #310 was one such regression, closed June 2026). Set OPENAI_SERVED_MODEL_NAME_OVERRIDE or pin the tagworker-vllm owner if a new release regresses
Jobs sit IN_QUEUE while workers idle, or the endpoint scaled itself down/health counters; SDK version in the image; worker logsrunpod-python 1.7.11 to 1.10.0 could stop workers pulling jobs on volume endpoints: upgrade to 1.10.1+. Repeated unhealthy workers trigger scale-down: fix the crash firstEngineering only if the upgrade does not fix it
Provisioning script will break on 15 NovemberGrep for rest.runpod.io/v1, flat bodies, bare-array parsing, /billing/endpointsREST v1 retirement. Job API calls on api.runpod.ai/v2 are unaffected; the billing path is the easy one to missProduct if they rely on Pod reset, which has no v2 action
Billed twice for long jobs; tax on credit purchasesJob IDs, execution time against executionTimeout, retries; billing countryLong jobs retried after timeout re-run without idempotency. Set timeout and TTL to the real job; dedupe in the handler. Tax depends on billing country; a valid tax ID may exempt a businessBilling team for a credit review, with job IDs attached

The demo's endpoint doctor and log triage tabs implement the first checks for these and eight more log signatures.

07

REST v1 retirement, 15 November 2026

The most time-boxed work an APAC FDE inherits. From Runpod's migration guide and the v2 OpenAPI schema.

Changev1 to v2What breaks if missed
Base URLrest.runpod.io/v1 to api.runpod.io/v2Every call, on the retirement date
Serverless billing path/billing/endpoints to /v2/billing/serverless/v2/billing/endpoints exists and returns Public Endpoint billing, so a naive rename returns the wrong data without an error
Pod lifecycle/pods/{id}/stop to POST /v2/pods/{id}/action with {"action":"stop"}reset has no v2 action; nightly reset jobs need a decision
Create bodiesFlat fields to nested gpu, workers, scaling, mounts; endpoint type required400 errors on create
List responsesBare arrays to {"pods":[...]} style wrappersLoops over the response iterate keys or fail
Errors{"message"} to RFC 9457 title, status, detailError handlers raise their own exceptions
GPU selection on PATCHgpu.pools and excludedTypes are replaced togetherSending pools alone clears exclusions and can land workers on unwanted GPUs

The v2 catalog endpoints need an API key, while today's GraphQL catalog query does not. Price trackers and scripts that read prices without a key will need one after GraphQL retires.

08

Highest-impact moves

1. APAC v2 migration sweep

List APAC accounts with v1 traffic, run each integration through a checker, and book fixes before 15 November. Protects spend at no acquisition cost.

2. Capacity-aware PoC template

Every APAC PoC ships with multiple GPU pools, no datacenter pin unless required, and one active worker on demo days. Removes the most common reason a PoC fails in front of a CTO.

3. APAC placement guide

One page on where to run from each APAC city, how data gets there without an APAC S3 endpoint, and when residency rules force a choice. Reused in every deal.

4. Serverless or Pod cost review

For accounts above a spend line, compare busy share with the break-even monthly. Moving the right workloads to active workers or reserved capacity lowers their bill and raises commitment.

5. Ticket signatures into docs

Second identical ticket becomes a KB article; fifth becomes a docs change or a product request with spend attached.

6. Capacity signal to supply

A monthly note to supply: APAC demand by GPU type and datacenter, lost or delayed deals, and the residency asks behind them.

09

First 90 days

WindowDoMeasured by
Days 1 to 30Shadow L2 and the AE; run every product end to end on my own account; v1 sweep for APAC accounts; take APAC-hours escalationsEvery APAC v1 integration identified, owners booked before 15 November
Days 31 to 60Own PoCs for the AE's pipeline; ship the PoC template and placement guide; first KB articles from repeat ticketsPoCs run against day-0 criteria; time to first PoC result
Days 61 to 90First production handovers; monthly feedback list to product; first APAC webinar or partner workshopOnboarded spend from PoCs; escalation time to resolution in APAC hours
10

Method & sources

Public sources only, read on 5 October 2026 (public job posting, 2026). Live figures come from Runpod's public GraphQL API. The demo is my own prototype built for this application.

Company: Series A post · $120M ARR note · TechCrunch

Pricing: runpod.io/pricing · Serverless pricing · Public Endpoints models

Serverless: operations · worker states · troubleshooting

vLLM: environment variables · worker-vllm #310 · runpod-python #402

API: Migrate from API v1 · v2 OpenAPI schema · release notes

Storage: S3-compatible API · global volumes

Status: uptime.runpod.io

Competitors: CoreWeave · Lambda · Modal · Together · Nebius · YTL AI

Model configs: Hugging Face config.json for each preset in the demo

Demo: runpod-fde.leverlabs.workers.dev

Independent homework for the Runpod Forward Deployed Engineer APAC role · 2026 · edwardtay.com