Data Software Engineer
Software Engineering
Remote
About Catena
Catena is the universal API for fleet telematics data. We give insurers, factors, TMS providers, brokers, fintechs, and freight-tech platforms a single integration to access carrier vehicle data across every major ELD and telematics provider, replacing dozens of one-off connections. Carriers opt in through a quick, permissioned flow, and our customers immediately get standardized, real-time data on a clean, normalized layer.
About the Role
We're hiring a Data Software Engineer to own the pipelines that turn 248+ telematics providers into one coherent data product.
Every provider models the same truck differently: a location ping, an hours-of-service event, a fuel level, and a fault code all arrive in a different shape, on a different cadence, with different gaps and different definitions of "now." Our customers never see that. They see one schema, one webhook, one API. Closing that gap is the hardest and most valuable engineering problem at Catena, and it's the one you'll work on.
This is a hands-on role on a small team. You'll build ingestion and normalization pipelines across a long tail of providers, design the storage and query patterns behind our APIs and Snowflake share, and build the derived data products our customers buy: stop detection, geofence events, ETAs, HOS availability, and fuel transaction matching. You'll also work on the write side: Catena doesn't just read from telematics systems, we write back into them, and that two-way path is what our customers can't get anywhere else.
Responsibilities
-
Build and operate ingestion pipelines across our telematics provider integrations, streaming, polling, and batch - with the retry, rate-limiting, and backpressure behavior each provider demands
-
Own the normalization layer: map wildly inconsistent provider payloads into Catena's canonical vehicle, driver, HOS, and fuel models, and defend that schema as new providers are added
-
Design storage and query patterns for high-volume time-series vehicle data, balancing real-time API latency against analytical workloads
-
Ship customer-facing data products, stop and geofence detection, ETA computation, location aggregates, IFTA-grade mileage - from spec through production
-
Build and maintain write-back workflows that push fuel transactions, DVIR status, and dispatch data back into provider systems, with idempotency and loop protection
-
Instrument everything: data freshness, field completeness, provider health, and drift detection, so we catch a broken upstream before a customer does
-
Own data quality end to end. In our business a silently wrong location is worse than a missing one - a fuel-card issuer declines a real transaction, an insurer misprices a carrier
-
Work directly with customers' engineering teams when the data question is theirs, not ours
Qualifications
-
3+ years building production data pipelines or backend data services
-
Strong Python; comfortable owning services end to end, not just notebooks or DAGs
-
Solid SQL and relational data modeling: you can reason about partitioning, indexing, and query cost, not just correctness
-
Experience with streaming or event-driven ingestion, and with the operational reality of unreliable third-party APIs
-
Cloud-native on AWS
-
You've debugged a pipeline that was quietly producing wrong data, and you have opinions about how to prevent it happening again
-
Clear written communication, we're remote and async by default
Our Tech Stack
Python · PostgreSQL / AWS Aurora · Snowflake · AWS (S3, Lambda, Kinesis) · webhooks and REST APIs · GitHub Actions · TypeScript on the application layer
Preferred
-
Experience in telematics, IoT, logistics, fintech, or insurance data - anywhere the data has physical-world ground truth and someone makes a money decision on it
-
Time-series or geospatial data at scale: geofencing, map matching, trip and stop inference, H3 or similar indexing
-
Data-sharing and multi-tenant delivery patterns - Snowflake shares, per-tenant credential isolation, consent-scoped access
-
Experience building against many third-party APIs at once, where you control neither the schema nor the uptime
-
Working in a SOC 2 environment, or helping get a company there
Why Catena, Why Now
-
Early-stage startup: high ownership, real impact, and a seat at the table
-
Work directly with leadership and have your voice heard on product and API direction
-
Be the technical authority customers rely on, and help build the Sales Engineering function from the ground up
-
Remote-friendly environment
Catena has raised $8.25M from Shaper Capital, Floating Point, Plug and Play, Liquid 2, and leading industry angels. Customers are already using us in production including modern TMS platforms (Mastery, Optym), large brokerages, and mid-to-large direct carriers operating 100–1,000+ trucks. Many of them are also investors.
Compensation: Competitive salary, full benefits, equity and 401k match.
US only, no sponsorship available. Occasional travel for team and customer on-sites.