Higgsfield AI** is the fastest-growing generative AI company in history: $500M annual recurring revenue (ARR), 25M+ users worldwide, 6M+ daily generations, and among our clients are 390 Fortune 500 brands. We are at the absolute frontier of AI video and next-generation creative tools. Joining Higgsfield means becoming part of a team that is shaping the future of AI-native products, in a company that doesn't just move fast, but redefines the very concept of 'fast'.
About the Team
The Data Acquisition team at Higgsfield AI is responsible for all aspects of data collection for training our models. We find, ingest, and prepare hundreds of millions of media assets and petabytes of video, images, and audio, working closely with the Data Processing, Infrastructure, ML, and Legal teams. We are looking for a strong engineer to join the Data Acquisition team.
What You'll Do
- Lead engineering projects in data acquisition: web crawling, ingestion, large-scale media collection.
- Develop and operate highly scalable distributed systems that handle petabytes of data.
- Design and write production services in Python for high-throughput concurrent processing, separating CPU-bound, I/O-bound, and GPU-bound pipeline stages.
- Build fault-tolerant ingestion from external sources: rate limiting, retries with backoff, connection pooling, correct operation during source degradation.
- Ensure pipeline resumability and idempotency: checkpoints, content-based deduplication, recovery without full processing restarts.
- Develop and maintain data storage backend services: large-scale object storage and high-load relational metadata stores.
- Deploy solutions in Kubernetes: autoscaling by queue depth, correct behavior on spot instances, regular system health checks.
- Instrument and monitor services using Prometheus, Grafana, VictoriaLogs, and Sentry, analyze performance, and resolve bottlenecks.
- Collaborate with ML teams on the composition, coverage, and quality of datasets to meet training requirements.
- Liaise with the Legal team on data compliance and privacy matters.
Requirements
- 3+ years of industrial development experience.
- Strong Python and experience developing production services operating under continuous load.
- Deep expertise in distributed systems and large-scale data processing: queue brokers, at-least-once semantics, idempotency, backpressure.
- Experience processing large multimodal media datasets.
- Solid understanding of the network stack and working with high-load HTTP clients.
- Proficient in Kubernetes and Infrastructure-as-Code practices.
- Ability to independently learn unfamiliar systems, protocols, and third-party code.
Bonus Points
- Experience with large crawlers, self-hosted storage systems, or media processing pipelines.
- Willingness to try new approaches and technologies.
- Ability to manage multiple tasks and adapt to changing priorities.
What We Offer
- Competitive base salary in USD — depending on your experience, skills, and role level.
- Stock Options: participation in the company's stock option program — an opportunity to share in Higgsfield's long-term growth.
- Relocation support to Almaty for candidates moving from another city or country.
- A dynamic, highly collaborative environment where you work directly with experienced leaders and have a real impact on the product and company.
- Opportunities for professional growth, responsibility, and career development as the company scales.
- Equipment, meals, transportation, and other office benefits covered by the company.