Start with the result, its deadline and the correctness rules before choosing a big-data tool. A large dataset does not automatically require a distributed processing framework.
Suppose the business needs yesterday's revenue by merchant each morning. That differs from detecting a suspicious payment before approving it. Both may involve many records, but one allows scheduled processing while the other needs a timely decision.
1. Define what must be processed
Clarify how much data exists, how quickly new data arrives, which fields are needed and how long data must be retained. For the revenue report, agree whether refunds change the original day or the day of the refund. Without that rule, a faster pipeline can still produce the wrong total.
