DineIQ Analytics – TechWiz Data Science Arena | Team NN-Azm (Hunain, Anas, Ishraq, Ammar)
1. The Business Problem A restaurant chain produces data every minute: orders, order lines, prices, ratings, promotions, stock consumption and food waste. Most managers still look at this through spreadsheets and monthly sales totals. That approach answers "what sold the most?" but it cannot answer the questions that decide profit: Which dish sells hundreds of plates a week but loses money on…
1. DineIQ Analytics is a web-based restaurant intelligence platform that uses Apache Spark, PySpark, and Spark SQL to process and analyze restaurant data. Two independent data science pipelines tackle the same problem using different programming languages and tools, with their outputs compared to ensure reliability.
2. Traditional restaurant reporting has three major weaknesses: it treats high sales volume as success regardless of profit, it is backward-looking, and it lacks verifiability. Restaurant data contains hidden relationships between multiple factors, making manual analysis impractical.
3. DineIQ Analytics classifies menu items into four categories - Profit Driver, Volume Driver, Hidden Opportunity, and Low Performer - using a combination of indicators. Evidence-based recommendations are provided for each suggestion, with a priority level assigned. The platform also generates RFM-based customer segments to identify high-value and at-risk customers.
4. The platform's big data architecture consists of multiple layers, including data generation, raw storage, Spark ingestion, data quality assessment, cleaning, Spark SQL integration, feature engineering, Parquet storage, prediction models, a recommendation engine, a database, and a web dashboard. The two ML pipelines are kept separate after the shared cleaned data to avoid one influencing the other's predictions.
5. The dataset for DineIQ Analytics was generated manually to mimic real business behavior. It includes tables like Customers, Orders, Order_Items, Menu_Items, Menu_Categories, Restaurants, Pricing_History, Promotions, Ratings, Inventory, Wastage, and more. The dataset contains millions of order lines, thousands of customers, and various attributes, such as price history, promotions, ratings, inventory, and wastage.
6. Before analysis, the platform conducts a data quality assessment, checking for missing values, duplicate records, invalid data, and inconsistencies. Cleaning follows documented rules, and every decision made during the process is recorded. Quarantining bad rows allows for review instead of silently deleting them.
7. DineIQ Analytics utilizes Apache Spark, PySpark, and Spark SQL for ingestion, integration, data cleaning, feature engineering, and more. Spark SQL is used for complex aggregations, such as revenue by location and month, as well as wastage by category.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.