STAFF DATA ENGINEER
Jobgether
Full-timesenior
Sign in to applyFree account, takes a minute.
Job description
Accountabilities
• Design, build, and own near real-time data pipelines using CDC, streaming ingestion, and event-driven architectures to power the platform’s core data flows.
• Evaluate, implement, and maintain vector database infrastructure and embedding pipelines supporting semantic search, retrieval-augmented generation, AI agents, and other AI-enabled applications.
• Build scalable ELT and ETL pipelines that ingest data from internal platforms, financial systems, CRM, HRIS, and other sources for both real-time and batch use cases.
• Partner with senior data leadership to architect the warehouse or lakehouse as the supporting system of record beneath streaming and AI infrastructure.
• Develop lightweight transformation layers using technologies such as dbt to help Analytics Engineering teams turn raw data into reliable, business-ready datasets and metrics.
• Own data pipeline reliability and observability, including monitoring, automated failure alerting, lineage tracking, and operational troubleshooting across streaming and batch environments.
• Establish the technical foundation for self-service and AI-powered reporting while partnering with BI, Product, and Engineering teams on executive and departmental reporting needs.
• Implement and maintain data governance practices covering documentation, access controls, lineage, and data quality standards.
• Collaborate with Finance, Marketing, Customer Success, Operations, and other stakeholders to translate business requirements into reliable, low-latency data products.
• Leverage AI-augmented development tools such as Claude Code or comparable solutions to accelerate engineering, testing, documentation, and workflow automation.
• Contribute to the evolution of the broader data architecture and establish scalable engineering practices as the platform grows.
Requirements
• 7+ years of hands-on data engineering experience, including substantial depth in streaming and event-driven architectures rather than exclusively batch processing.
• Proven experience designing and building production-grade near real-time pipelines from the ground up using technologies such as Kafka, Kinesis, Flink, Debezium, CDC, or similar tools.
• Hands-on production experience with vector databases and embedding infrastructure, such as Pinecone, Weaviate, pgvector, Milvus, Zilliz, or comparable technologies.
• Experience developing embedding strategies and chunking approaches for retrieval and AI-powered use cases is strongly valued.
• Advanced proficiency in SQL and Python.
• Working knowledge of cloud warehouse or lakehouse platforms such as Snowflake, BigQuery, or Databricks, as well as dbt.
• Proven experience building or materially contributing to an end-to-end production data environment, ideally as an early, founding, or highly autonomous data engineering hire.
• Familiarity with B2B SaaS data models, including customer lifecycle, sales pipeline, conversion, recurring revenue, ARR, CAC, and churn concepts.
• Working knowledge of BI and reporting tools such as Looker, Tableau, Power BI, or similar platforms.
• Familiarity with ELT/ETL tools such as Fivetran or Airbyte for batch data integration.
• Strong understanding of data observability and reliability practices, including automated alerting, lineage tracking, monitoring, and failure recovery.
• Experience using AI-augmented engineering and development tools such as Claude, Copilot, or similar technologies.
• Ability to work independently, manage complex multi-stakeholder initiatives, and make sound technical decisions in a lean, fast-moving environment.
• Comfortable operating with ambiguity and building new infrastructure, processes, and standards where limited existing foundations are in place.
Benefits
• Competitive compensation package based on skills, experience, qualifications, and work location.
• Health benefits for eligible U.S.-based employees.
• Flexible paid time off.
• Parental leave.
• Fertility and adoption assistance.
• 401(k) benefits.
• Educational reimbursement.
• Opportunities to work with modern streaming, AI, vector database, and data platform technologies.
• Meaningful ownership over the development of a new end-to-end data platform.
• Support for reasonable accommodations throughout the application and interview process.
• A collaborative environment focused on building scalable technology and continuously improving data capabilities.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Skills
KafkaKinesisFlinkDebeziumCDCPineconeWeaviatepgvectorMilvusZillizSnowflakeBigQueryDatabricksdbtFivetranAirbyteSQLPythonLookerTableauPower BIClaudeCopilot