YOUR ROLE
Data Scientist – Classical Machine Learning (Equipment Manufacturing)
Experince : 7 years and above
About the Role
We are seeking a Data Scientist with strong expertise in classical machine learning and predictive analytics to develop custom models for equipment manufacturing and industrial environments. The ideal candidate will have hands-on experience building machine learning models from the ground up, selecting the most suitable algorithms based on business needs, and transforming manufacturing data into actionable insights.
This role focuses on addressing real-world industrial challenges, including equipment failure prediction, quality forecasting, demand forecasting, process optimization, predictive maintenance, and production efficiency enhancement.
Key Responsibilities
- Develop predictive models using classical machine learning techniques.
- Design, train, validate, and deploy machine learning models using manufacturing and operational data.
- Analyze structured and time-series data from production equipment, sensors, ERP systems, MES platforms, and quality management systems.
- Perform feature engineering, data exploration, and statistical analysis.
- Build scalable data preprocessing, model training, and validation pipelines.
- Evaluate multiple algorithms and identify the most effective solutions based on business objectives and performance metrics.
- Interpret model outputs and communicate insights to engineering, operations, and business stakeholders.
- Collaborate with manufacturing, process engineering, quality assurance, and maintenance teams to understand operational challenges and identify opportunities for data-driven improvements.
- Deploy machine learning models into production environments and monitor model performance.
- Continuously improve model accuracy, reliability, and scalability.
Required Technical Skills
Machine Learning
Strong understanding and practical experience with classical machine learning algorithms, including:
- Linear Regression
- Logistic Regression
- Decision Trees
- Random Forest
- Gradient Boosting
- XGBoost
- LightGBM
- CatBoost
- Support Vector Machines (SVM)
- K-Nearest Neighbors (KNN)
- Naïve Bayes
- Clustering Techniques (K-Means, DBSCAN, Hierarchical Clustering)
- Principal Component Analysis (PCA) and Dimensionality Reduction
- Time Series Forecasting (ARIMA, SARIMA, Prophet)
- Ensemble Learning Methods
- Anomaly Detection Techniques
YOUR PROFILE
Statistical Knowledge
- Hypothesis Testing
- Probability
- Experimental Design
- Statistical Modeling
- Regression Analysis
- Sampling Techniques
- Confidence Intervals
- Time Series Analysis
Programming
- Python
- SQL
Python libraries:
- pandas
- NumPy
- scikit-learn
- SciPy
- statsmodels
- XGBoost
- LightGBM
- CatBoost
- Matplotlib
- Seaborn
Data Engineering Skills
- Data Cleaning
- Feature Engineering
- Data Integration
- ETL Pipelines
- Handling Missing Values
- Data Validation
Manufacturing Domain Knowledge (Preferred) Experience with:
- Industrial Manufacturing
- Equipment Manufacturing
- Automotive
- Heavy Engineering
- Process Manufacturing
- Factory Automation
Understanding of:
- Production KPIs
- OEE
- Downtime Analysis
- Root Cause Analysis
- MES
- ERP
- SCADA
- PLC Data
- Sensor Data
- IIoT
Model Development Expectations Candidates should demonstrate the ability to:
- Select appropriate algorithms based on business problems.
- Build machine learning models from scratch using Python.
- Engineer meaningful features from manufacturing datasets.
- Optimize hyperparameters.
- Evaluate models using appropriate metrics.
- Explain model decisions using interpretable techniques.
- Deploy models into production.
Experience with AutoML tools alone is not sufficient.
Required Qualifications
- Strong problem-solving and analytical skills.
- Experience working with structured industrial datasets.
- Excellent communication and stakeholder management skills.
Preferred Qualifications
- Experience in predictive maintenance projects.
- Experience with manufacturing analytics.
- Knowledge of MLOps practices.
- Familiarity with Docker and cloud platforms (AWS, Azure, or GCP).
- Exposure to edge analytics or IoT-based machine learning.
- Experience integrating ML models into enterprise applications.
Nice to Have
- Knowledge of optimization techniques.
- Survival Analysis and Remaining Useful Life (RUL) modeling.
- Reinforcement Learning (basic understanding).
- Digital Twin concepts.
- Knowledge of reliability engineering.
- Experience working with streaming data.
Success Measures Within the first 6-12 months, the successful candidate should be able to:
- Develop production-ready predictive models for equipment health, quality, or process optimization.
- Improve prediction accuracy through feature engineering and model tuning.
- Collaborate effectively with manufacturing and engineering teams to deliver measurable business outcomes.
- Deploy and monitor machine learning solutions that support operational decision-making.