A data-driven exploration into the patterns, causes, and insights behind train delays in India, using over 2.8 million data points collected across 3 months and 3,100 trains. From fog-induced disruptions to station-level bottlenecks β discover how data reveals the untold stories of our railways.
- π¦ Dataset: 2.8M+ rows of real-time train status logs
- π οΈ Tools Used: Python, Pandas, Matplotlib, Seaborn, Frida, BurpSuite, ADB, Linux shell scripting
- π Scope:
- Delay variation by station, distance, time of day, weekday, and season
- Impact of fog and winter on punctuality
- Identification of high-delay stations and resilient nodes
- π Output: A series of clean, publication-ready visualizations + LaTeX report
| π Insight | π§ What We Found |
|---|---|
| Stations with extreme delays | Gumgaon delays π¨, Virinchipuram early π |
| Delay vs Distance from Origin | Longer routes β higher delays |
| Delay vs Time of Day | Peak: 12β16 hrs β° |
| Delay vs Day of Week | Mondays worst, Fridays best |
| Delay vs Month | Jan: π₯ delays, March: β smooth rides |
| Fog Impact on Trains | 20,000+ 2hr+ delays in Jan due to fog π«οΈ |
- π Prioritize delay-prone stations for audits
- β±οΈ Reschedule high-importance trains to off-peak hours
- βοΈ Plan fog-safe routing