Principal Software Engineer, Platform Services (PMTS)

Airkit
Airkit

Software Engineering

San Francisco, CA, USA

Posted on Sep 12, 2026

Description

Summary

Design, build, and maintain the backend services, APIs, data-platform automation, and AI platform / agent systems behind our data products. This is a backend-heavy role: distributed Python services, streaming. LLM/agent runtimes on AWS, RAG pipelines, and the dbt/Airflow automation that powers a Snowflake-based lakehouse.


You will work inside the Developer Experience team, whose mandate is to build the tooling, "paved paths," and automation that let internal data teams ship discoverable, governed, and consumable data products quickly and safely.

About the platform you'll work on

  • Lakehouse & warehouse — Snowflake with Apache Iceberg tables in a medallion architecture (Bronze → Silver → Gold → Semantic Views).

  • Transformationdbt (Cloud on Fusion) with data contracts, tag-driven governance, and automated project-evaluation checks.

  • OrchestrationAirflow on AWS MWAA; DAGs and schedules authored and deployed through platform tooling and CI (metadata driven).

  • Developer tooling — an internal CLI and reusable utilities (dbt, Iceberg operations, Snowflake load, REST API based ingestions, etc.) that abstract the platform for data teams.

  • Data quality & governanceMonte Carlo for data observability (monitors as code) and centralized governance (tagging, lineage, PII classification, cataloging).

  • AI layer — Agent/MCP/Skill deployment framework, RAG, semantic router / broker agents, over our data products and catalog metadata; conversational agents in Slack;
    model access via AWS Bedrock / AgentCore, centralized LLM Gateway, and Snowflake Cortex.

  • Cloud & delivery — AWS (ECS Fargate, Lambda, API Gateway, DynamoDB, S3, IAM, etc.), Terraform IaC, GitHub Actions CI/CD, containerized builds, structured logging, metrics, and tracing.


​Responsibilities

  • Backend services

  • Design, build, and maintain backend services in Python / Go to support our Data Platform Services — Agent runtimes on AWS Bedrock AgentCore / ECS Fargate, and event-driven Lambda handlers for our AI Platform Services

  • Design and maintain our abstraction layers for for our data domain customers via custom developed utilities

  • Maintain and service our infrastructure using IaC (Terraform) for our AWS accounts

  • AI / agents

  • Build RAG pipelines and tool-calling AI agents for our data products: retrieval, orchestration, grounding/citations, and evaluation.

  • Integrate model providers with our LLM Gateway with response streaming, prompt caching, and structured outputs in an auditable way (OTeL)

  • Stand up eval harnesses, guardrails, and tracing so agent quality, latency, and cost are measurable and regressions are caught before release.

  • Data-platform automation

  • Build and maintain data-platform automation: dbt services, MWAA/Airflow orchestration, and tooling that makes data products discoverable, governed, and consumable.

  • Extend the platform CLI and shared utilities that data teams use as their day-to-day interface to the platform.

  • Identity & security

  • Implement auth and identity: OAuth/OIDC flows (per-user 3LO, token vaulting, session binding), least-privilege IAM, and secrets management.

  • Enforce multi-tenant isolation and per-user identity/RBAC across services and agents.

  • Reliability & delivery

  • Own service reliability and delivery: Terraform, GitHub Actions CI/CD, container builds, structured logging, metrics/tracing, alerting, and cost controls.

  • Set technical direction: system and API design, code review, and mentoring.

Required skills

Backend engineering

  • 15+ years designing, building, and maintaining production backend services at scale

  • Expert-level Python / Go / Java for server-side development; solid grasp of relevant frameworks and the WSGI/ASGI model.

  • Service & API design: REST (and/or gRPC), request/response and streaming patterns, pagination, versioning, idempotency, and backward-compatible contracts.

  • Data layer: SQL and data modeling, query optimization and indexing, transactions, connection pooling; relational, warehouse (Snowflake), and NoSQL/key-value (DynamoDB) stores.

  • Server-side patterns: caching strategies, background jobs/workers, queues and event-driven processing, rate limiting, retries/backoff, and timeouts.

  • Performance & reliability: profiling, load handling, latency/throughput trade-offs, graceful degradation, and designing for failure.

  • Observability: structured logging, metrics, distributed tracing, and debugging live production issues.

Software engineering fundamentals

  • OOP (required): encapsulation, abstraction, inheritance, composition, polymorphism; SOLID principles; design patterns applied pragmatically; strong domain modeling.

  • Solid data structures & algorithms; ability to reason about time/space complexity.

  • Concurrency & async programming (async/await, threading, event loops) and their failure modes.

  • Testing (unit, integration, end-to-end) and testable design; Git and PR-based workflows; disciplined code review.

Cloud & infrastructure

  • Production AWS: ECS/containers, Lambda, IAM, API Gateway, DynamoDB.

  • Infrastructure as code with Terraform; CI/CD (GitHub Actions or equivalent) and container builds.

Distributed systems

  • Building services that are horizontally scalable, resilient, and loosely coupled; handling consistency, retries, idempotency, and partial failure.

Security

  • OAuth/OIDC, authn/authz, token handling, least-privilege access, multi-tenant isolation, secrets management.

AI / LLM engineering

  • Building LLM applications in production (not research).

  • RAG: chunking, embeddings, vector search, hybrid search, reranking, grounding/citations, context-window management, retrieval evaluation.

  • Agents: prompt engineering, tool use/function calling, structured outputs, single- and multi-step orchestration, prompt caching.

  • Integration with model providers — Anthropic/Claude, AWS Bedrock/AgentCore, Snowflake Cortex — and response streaming.

  • AI quality & ops: eval harnesses, guardrails, tracing/observability, token/latency/cost optimization.

  • AI security: prompt injection, data exfiltration, PII handling, per-user identity/RBAC enforcement.

Nice to have

  • dbt, Airflow/MWAA, Snowflake hands-on experience.

  • MCP (Model Context Protocol) and/or applied RAG.

  • Slack platform (Bolt, Socket Mode, Block Kit) or other real-time/conversational backends.

  • Data observability (Monte Carlo) or data-quality/governance tooling.

  • Fine-tuning/adaptation, semantic caching, or model routing/fallback.

  • Experience adding AI capabilities to existing production systems.

  • Iceberg / open table formats and lakehouse patterns.

  • Tableau or other BI integration.

Pursuant to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Salesforce will consider for employment qualified applicants with arrest and conviction records.

In the United States, compensation offered will be determined by factors such as location, job level, job-related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity, and benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link: https://www.salesforcebenefits.com.