by Anand Rajaraman, Jeffrey D. Ullman, Jure Leskovec
The book is based on Stanford Computer Science course CS246: Mining Massive Datasets (and CS345A: Data Mining).