COMP 465 Special Topics: Data Mining

Similar documents
Introduction to Data Mining S L I D E S B Y : S H R E E J A S W A L

Knowledge Discovery and Data Mining

Data Mining. Chapter 1: Introduction. Adapted from materials by Jiawei Han, Micheline Kamber, and Jian Pei

Data Mining: Concepts and Techniques

CS423: Data Mining. Introduction. Jakramate Bootkrajang. Department of Computer Science Chiang Mai University

This tutorial has been prepared for computer science graduates to help them understand the basic-to-advanced concepts related to data mining.

CSE-4412: Data Mining

Winter Semester 2009/10 Free University of Bozen, Bolzano

Big Data Analytics The Data Mining process. Roger Bohn March. 2016

cse643, cse590 Data Mining Anita Wasilewska Computer Science Department Stony Brook University Stony Brook NY 11794

cse643 Data Mining Professor Anita Wasilewska Computer Science Department Stony Brook University

Summary of Last Class. Course Content. Chapter 1 Objectives

Data Mining. Yi-Cheng Chen ( 陳以錚 ) Dept. of Computer Science & Information Engineering, Tamkang University

CS377: Database Systems Data Warehouse and Data Mining. Li Xiong Department of Mathematics and Computer Science Emory University

Chapter 1, Introduction

CS590D: Data Mining Chris Clifton. What Is Data Mining?

Data Mining: Concepts and Techniques. (3 rd ed.) Chapter 1

DATA MINING II - 1DL460

Data warehouse and Data Mining

Question Bank. 4) It is the source of information later delivered to data marts.

DATA MINING II - 1DL460

Concepts and Techniques. Data Mining: Slides related to: University of Illinois at Urbana-Champaign

Applications and Trends in Data Mining

CSE5243 INTRO. TO DATA MINING

INTRODUCTION TO DATA MINING

D B M G Data Base and Data Mining Group of Politecnico di Torino

Big Data Analytics The Data Mining process. Roger Bohn March. 2017

What is Data Mining?

COMP90049 Knowledge Technologies

Data Mining Course Overview

Knowledge Discovery & Data Mining

Knowledge Discovery. URL - Spring 2018 CS - MIA 1/22

ISM 50 - Business Information Systems

Introduction to Data Mining and Data Analytics

Knowledge Discovery in Data Bases

Study on the Application Analysis and Future Development of Data Mining Technology

Knowledge Discovery. Javier Béjar URL - Spring 2019 CS - MIA

Data mining fundamentals

An Introduction to Data Mining BY:GAGAN DEEP KAUSHAL

Database and Knowledge-Base Systems: Data Mining. Martin Ester

International Journal of Advance Engineering and Research Development. A Survey on Data Mining Methods and its Applications

WKU-MIS-B10 Data Management: Warehousing, Analyzing, Mining, and Visualization. Management Information Systems

Lecture 18. Business Intelligence and Data Warehousing. 1:M Normalization. M:M Normalization 11/1/2017. Topics Covered

GUJARAT TECHNOLOGICAL UNIVERSITY MASTER OF COMPUTER APPLICATIONS (MCA) Semester: IV

Chapter 28. Outline. Definitions of Data Mining. Data Mining Concepts

CISC 4631 Data Mining Lecture 01:

Computers Are Your Future

Data Mining Concept. References. Why Mine Data? Commercial Viewpoint. Why Mine Data? Scientific Viewpoint

Thanks to the advances of data processing technologies, a lot of data can be collected and stored in databases efficiently New challenges: with a

Data Mining. Introduction. Piotr Paszek. (Piotr Paszek) Data Mining DM KDD 1 / 44

1. Inroduction to Data Mininig

Data Mining. Prof. Jiawei Han of UIUC

Table Of Contents: xix Foreword to Second Edition

UNIT-IV DATA MINING ALGORITHMS. March 14, 2012 Prof. Asha Ambhaikar

Data Mining. Ryan Benton Center for Advanced Computer Studies University of Louisiana at Lafayette Lafayette, La., USA.

Introduction. to Data Mining. Introduction. Motivation: Necessity is the Mother of Invention. Motivation: Necessity is the Mother of Invention

Data Mining and Data Warehousing Introduction to Data Mining

Contents. Foreword to Second Edition. Acknowledgments About the Authors

INSTITUTE OF AERONAUTICAL ENGINEERING (Autonomous) Dundigal, Hyderabad

Data Mining Concepts

Design and Realization of Data Mining System based on Web HE Defu1, a

DATA WAREHOUSING AND MINING UNIT-V TWO MARK QUESTIONS WITH ANSWERS

Data Mining Technology Based on Bayesian Network Structure Applied in Learning

Data Mining. Vera Goebel. Department of Informatics, University of Oslo

Introduction. IST557 Data Mining: Techniques and Applications. Jessie Li, Penn State University

Foundation of Data Mining: Introduction

The 2018 (14th) International Conference on Data Science (ICDATA)

Chapter 3: Data Mining:

A SURVEY OF DATA MINING & ITS APPLICATIONS

Data mining overview. Data Mining. Data mining overview. Data mining overview. Data mining overview. Data mining overview 3/24/2014

CS246: Mining Massive Datasets Jure Leskovec, Stanford University.

CS 412 Intro. to Data Mining

INTRODUCTION TO BIG DATA, DATA MINING, AND MACHINE LEARNING

Knowledge Modelling and Management. Part B (9)

Dynamic Data in terms of Data Mining Streams

Data Mining Concepts. Duen Horng (Polo) Chau Assistant Professor Associate Director, MS Analytics Georgia Tech

Statistical Learning and Data Mining CS 363D/ SSC 358

Introduction to Data Mining

Data Mining: Approach Towards The Accuracy Using Teradata!

3 Data, Data Mining. Chengkai Li

Data Mining & Machine Learning

CS249: ADVANCED DATA MINING

Overview of Web Mining Techniques and its Application towards Web

Data Mining. Introduction. Hamid Beigy. Sharif University of Technology. Fall 1395

Parametric Comparisons of Classification Techniques in Data Mining Applications

DATA MINING AND DATABASE TECHNOLOGY (WEB MINING, TEXT MINING, SENTIMENTAL ANALYSIS FOR SOCIAL MEDIA, TOOLS, TECHNIQUES, METHODS,

Chapter 1 Introduction to Data Mining

Introduction to Data Mining

Data Mining. Introduction. Hamid Beigy. Sharif University of Technology. Fall 1394

Volume 3, Issue 9, September 2015 International Journal of Advance Research in Computer Science and Management Studies

Now, Data Mining Is Within Your Reach

Business Data Analytics

9. Conclusions. 9.1 Definition KDD

An Introduction to Data Mining in Institutional Research. Dr. Thulasi Kumar Director of Institutional Research University of Northern Iowa

CT75 (ALCCS) DATA WAREHOUSING AND DATA MINING JUN

INTRODUCTION... 2 FEATURES OF DARWIN... 4 SPECIAL FEATURES OF DARWIN LATEST FEATURES OF DARWIN STRENGTHS & LIMITATIONS OF DARWIN...

Data Mining and Warehousing

DEPARTMENT OF COMPUTER SCIENCE AND ENGINEERING

Defining a Data Mining Task. CSE3212 Data Mining. What to be mined? Or the Approaches. Task-relevant Data. Estimation.

Microsoft Exam

Transcription:

COMP 465 Special Topics: Data Mining Introduction & Course Overview 1

Course Page & Class Schedule http://cs.rhodes.edu/welshc/comp465_s15/ What s there? Course info Course schedule Lecture media (slides, handouts, etc) Assignments (reading & things to hand in) Project information (coming soon ) Presentation information (coming soon ) 2

Course Topics Data acquisition, data warehousing Data mining process Association Rule Mining Introduction to WEKA Classification Clustering Neural Networks Big Data 3

Course Grades 30% In-Class Exercises, Online quizzes 10% Paper Presentation 25% Final Project 15% Midterm Thursday, March 5 th in class 20% Final 4

What Is Data Mining? Data mining (knowledge discovery from data) Extraction of interesting (non-trivial, implicit, previously unknown and potentially useful) patterns or knowledge from huge amount of data Data mining: a misnomer? Alternative names Knowledge discovery (mining) in databases (KDD), knowledge extraction, data/pattern analysis, data archeology, data dredging, information harvesting, business intelligence, etc. Watch out: Is everything data mining? Simple search and query processing (Deductive) expert systems 5

What is Data Mining? Real Example from the NBA Play-by-play information recorded by teams Who is on the court Who shoots Results Coaches want to know what works best Plays that work well against a given team Good/bad player matchups Advanced Scout (from IBM Research) is a data mining tool to answer these questions Starks+Houston+ Ward playing Overall 0 20 40 60 Shooting Percentage 6

What is Data Mining? 7

Integration of Multiple Technologies Machine Learning Artificial Intelligence Database Management Statistics Algorithms Visualization Data Mining 8

Knowledge Discovery (KDD) Process This is a view from typical database systems and data warehousing communities Data mining plays an essential role in the knowledge discovery process Data Mining Pattern Evaluation Task-relevant Data Data Warehouse Selection Data Cleaning Data Integration Databases

Example: A Web Mining Framework Web mining usually involves Data cleaning Data integration from multiple sources Warehousing the data Data cube construction Data selection for data mining Data mining Presentation of the mining results Patterns and knowledge to be used or stored into knowledge-base 10

Data Mining in Business Intelligence Increasing potential to support business decisions Decision Making Data Presentation Visualization Techniques Data Mining Information Discovery End User Business Analyst Data Analyst Data Exploration Statistical Summary, Querying, and Reporting Data Preprocessing/Integration, Data Warehouses Data Sources Paper, Files, Web documents, Scientific experiments, Database Systems DBA

KDD Process: A Typical View from ML and Statistics Input Data Data Pre- Processing Data Mining Post- Processing Data integration Normalization Feature selection Dimension reduction Pattern discovery Association & correlation Classification Clustering Outlier analysis Pattern evaluation Pattern selection Pattern interpretation Pattern visualization This is a view from typical machine learning and statistics communities

Why Data Mining? The Explosive Growth of Data: from terabytes to petabytes Data collection and data availability Automated data collection tools, database systems, Web, computerized society Major sources of abundant data Business: Web, e-commerce, transactions, stocks, Science: Remote sensing, bioinformatics, scientific simulation, Society and everyone: news, digital cameras, YouTube We are drowning in data, but starving for knowledge! Necessity is the mother of invention Data mining Automated analysis of massive data sets 13

Why Data Mining? Potential Applications Data analysis and decision support Market analysis and management Target marketing, customer relationship management (CRM), market basket analysis, cross selling, market segmentation Risk analysis and management Forecasting, customer retention, improved underwriting, quality control, competitive analysis Fraud detection and detection of unusual patterns (outliers) Other Applications Text mining (news group, email, documents) and Web mining Stream data mining DNA and bio-data analysis 14

Market Analysis and Management Where does the data come from? Credit card transactions, loyalty cards, discount coupons, customer complaint calls, plus (public) lifestyle studies Target marketing Find clusters of model customers who share the same characteristics: interest, income level, spending habits, etc. Determine customer purchasing patterns over time Cross-market analysis Associations/co-relations between product sales, & prediction based on such association Customer profiling What types of customers buy what products (clustering or classification) Customer requirement analysis identifying the best products for different customers predict what factors will attract new customers Provision of summary information multidimensional summary reports statistical summary information (data central tendency and variation) 15

Corporate Analysis & Risk Management Finance planning and asset evaluation cash flow analysis and prediction contingent claim analysis to evaluate assets cross-sectional and time series analysis (financial-ratio, trend analysis, etc.) Resource planning summarize and compare the resources and spending Competition monitor competitors and market directions group customers into classes and a class-based pricing procedure set pricing strategy in a highly competitive market 16

Fraud Detection & Mining Unusual Patterns Approaches: Clustering & model construction for frauds, outlier analysis Applications: Health care, retail, credit card service, telecomm. Auto insurance: ring of collisions Money laundering: suspicious monetary transactions Medical insurance Professional patients, ring of doctors, and ring of references Unnecessary or correlated screening tests Telecommunications: phone-call fraud Phone call model: destination of the call, duration, time of day or week. Analyze patterns that deviate from an expected norm Retail industry Analysts estimate that 38% of retail shrink is due to dishonest employees Anti-terrorism 17

Other Applications Sports IBM Advanced Scout analyzed NBA game statistics (shots blocked, assists, and fouls) to gain competitive advantage for New York Knicks and Miami Heat Astronomy JPL and the Palomar Observatory discovered 22 quasars with the help of data mining Internet Web Surf-Aid IBM Surf-Aid applies data mining algorithms to Web access logs for market-related pages to discover customer preference and behavior pages, analyzing effectiveness of Web marketing, improving Web site organization, etc. 18

What Can Data Mining Do? Cluster Classify Categorical, Regression Summarize Summary statistics, Summary rules Link Analysis / Model Dependencies Association rules Sequence analysis Time-series analysis, Sequential associations Detect Deviations 19

Data Mining Complications Volume of Data Clever algorithms needed for reasonable performance Interest measures How do we ensure algorithms select interesting results? Knowledge Discovery Process skill required How to select tool, prepare data? Data Quality How do we interpret results in light of low quality data? Data Source Heterogeneity How do we combine data from multiple sources? 20

Major Issues in Data Mining (1) Mining Methodology Mining various and new kinds of knowledge Mining knowledge in multi-dimensional space Data mining: An interdisciplinary effort Boosting the power of discovery in a networked environment Handling noise, uncertainty, and incompleteness of data Pattern evaluation and pattern- or constraint-guided mining User Interaction Interactive mining Incorporation of background knowledge Presentation and visualization of data mining results 21

Major Issues in Data Mining (2) Efficiency and Scalability Efficiency and scalability of data mining algorithms Parallel, distributed, stream, and incremental mining methods Diversity of data types Handling complex types of data Mining dynamic, networked, and global data repositories Data mining and society Social impacts of data mining Privacy-preserving data mining Invisible data mining 22

Next Time Overview: Data mining tasks - Clustering, Classification, Rule learning, etc. Read Chapter 1 in Han Book 23