
Remote opportunity at
Gather AISoftware Development Engineer II – Data Engineer
Gather AI operates in the robotics and supply chain technology sector, developing a platform that uses autonomous drones and existing equipment to collect real-time data and…
Career Tools
About This Role
Gather AI operates in the robotics and supply chain technology sector, developing a platform that uses autonomous drones and existing equipment to collect real-time data and digitize warehouse workflows. The organization is forming a new Full Stack group within its Cloud Services division to construct a central data foundation, analytical warehouse, and semantic model from the ground up. As a Software Development Engineer II focused on Data Engineering, the selected…
Job Description
Gather AI operates in the robotics and supply chain technology sector, developing a platform that uses autonomous drones and existing equipment to collect real-time data and digitize warehouse workflows. The organization is forming a new Full Stack group within its Cloud Services division to construct a central data foundation, analytical warehouse, and semantic model from the ground up.
As a Software Development Engineer II focused on Data Engineering, the selected professional will help establish this foundational data architecture. Rather than maintaining an existing legacy platform, this position involves building new data foundations, shifting analytics away from the production PostgreSQL database, and developing shared dbt models alongside senior engineers. Initial work focuses on the Drone product before expanding to other company products.
This is a fully remote position intended for candidates based in India. Successful execution requires operating across multiple time zones, working independently, and utilizing clear written communication skills.
Responsibilities
- Develop extraction pipelines to move data from production PostgreSQL into the analytical warehouse using incremental loads
- Write and maintain shared dbt models including dimensions, serving tables, and metric building blocks
- Implement consistent metrics within the semantic layer across all dashboards
- Add data quality tests, freshness checks, alerts, and manage safe backfills
- Apply tenant separation and access control patterns to pipelines and models
- Maintain data lineage linking metrics back to original source records and drone images
- Collaborate with the integration team to validate incoming warehouse management system data
- Document models within the data catalog
- Deploy changes using continuous integration and delivery pipelines while participating in on-call rotations
- Expand data foundation support from drone products to Material Handling Equipment vision and 3D case counting products
Requirements
- Two to five years of experience building and operating production data pipelines
- Strong SQL proficiency including window functions, CTEs, joins, and a working understanding of dimensional modeling
- Hands-on experience building tested models using dbt, Snowflake Dynamic Tables, Databricks Lakeflow Declarative Pipelines, or equivalent tools
- Production pipeline coding experience in Python with testing and code reviews
- Experience running pipelines and managing failures, re-runs, and backfills in Airflow, Dagster, Databricks Lakeflow Jobs, Snowflake Tasks, or similar orchestration tools
- Experience loading operational database data into a warehouse or lakehouse like Snowflake or Databricks
- Production cloud experience on Azure or another major cloud provider involving object storage, Git, CI/CD, and Docker or Kubernetes
- Clear written and spoken English communication skills
Qualifications
- A degree in Computer Science or equivalent practical experience
- Experience with change data capture tools such as Debezium or Fivetran, or streaming technologies like Kafka or Event Hubs
- Familiarity with PySpark or Snowpark
- Experience with Terraform or infrastructure as code
- Familiarity with semantic layers like the dbt Semantic Layer or Cube, and data catalogs like Purview or DataHub
- Experience with multi-tenant data platforms or row-level security
- Experience working with image, video, or sensor data alongside structured records
- Domain experience in logistics, warehousing, or robotics
Core Skills
Frequently Asked Questions
Is this position remote?
Yes, this is a fully remote role on the company's India-based team.
What is the employment type?
The posting specifies that this is a full-time position.
What is the salary for this role?
The job posting does not specify a salary.
What experience level is required?
The role requires two to five years of experience building and running production data pipelines, alongside a degree in Computer Science or equivalent practical experience.
Sample Interview Questions
AI-generated questions tailored to this specific role — a preview of the full practice set.