About the Role
In this pivotal role, you will be responsible for building and managing the project's core data infrastructure. Your primary focus will be to develop a robust software platform for large-scale acquisition, processing, and management of complex, multimodal datasets (including EEG, EMG, and acoustics) from our hardware prototypes. This platform is foundational to all machine learning research and development activities. You will work closely with our research and machine learning teams to build data pipelines that facilitate feature engineering, dataset versioning, and model optimization, playing an essential part in advancing our innovative technology.
What You’ll Do
- Full-Stack Platform Development: Design, build, and manage a scalable, full-stack research data platform, encompassing backend services, robust databases (SQL/NoSQL), and intuitive frontend interfaces for data visualization and annotation.
- ML Pipeline Integration: Engineer and maintain robust data pipelines for continuous, multi-modal data acquisition (EEG, EMG, acoustics, video) from our hardware prototypes, directly supporting the entire machine learning lifecycle from data ingestion and feature extraction to model training and validation.
- ML Workflow Enablement: Collaborate closely with the machine learning team to understand their data requirements, streamline dataset preparation, and build tools that accelerate the experimentation and optimization of complex models.
- System Architecture & Integrity: Work with hardware and research teams to define data requirements and architect data flows, ensuring the highest level of data integrity and consistency across the project.
- Code Quality & Documentation: Produce clean, maintainable, and well-tested code, along with comprehensive documentation for all software, APIs, and data protocols.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Software Engineering, or a related field.
- Strong proficiency in Python and hands-on experience in full-stack web development.
- Proven experience in software engineering fundamentals, including designing and building scalable backend systems, developing REST APIs, and managing databases (SQL/NoSQL).
- A strong grasp of the end-to-end machine learning workflow, from data collection and preprocessing to model training and evaluation.
- Experience working with time-series data (e.g., EEG, EMG, audio, or other sensor data) in an ML context.
- Familiarity with common ML frameworks such as PyTorch or TensorFlow and version control with Git.
- A proactive, problem-solving mindset with the ability to work independently and manage tasks in a dynamic research environment.
Preferred Qualifications