中正大學課程大綱
Seminar in Data Mining and Big Data Analytics資料探勘與大數據分析專題研究
一、課程概述
國立中正大學課程大綱 115學年第1學期
教育學院教育研究所博士班-體育運動組
課程名稱(中文) 資料探勘與大數據分析專題研究
課程名稱(英文) Seminar on Data Mining and Big Data Analytics
授課教師 林晉榮 博士/教授
學分/時數 2 學分
授課方式 講授及實機操作
開課時段 星期二 上午 10:10-12:00(實際時段依系所排課為準)
上課地點 電腦教室(每人一機,需安裝 IBM SPSS Modeler)
先修科目或先備能力 無。本課程不要求任何程式設計或資訊背景,全程以圖形化介面操作。
課程屬性 博士班專題研究(Seminar)——以「專案分析(Project-based Analytics)」為主要學習形式

一、課程概述
本課程為博士班專題研究課程,以《資料探勘:人工智慧與機器學習發展-以 SPSS Modeler 為範例》為主要教材,並以「體育運動數據分析」作為全部應用實例之場域。課程設計不同於碩士班的工具操作導向,而以「專案分析(Project-based Analytics)」為主軸:學生自學期第一週起選定一個具研究價值的體育運動資料探勘題目,循 CRISP-DM(Cross-Industry Standard Process for Data Mining)六階段——商業理解、資料理解、資料準備、建模、評估、佈署——逐週推進,於學期末產出一份具投稿水準的專題分析報告。
考量修課學生為體育運動專業背景、多無資訊或程式設計基礎,本課程採「零程式(No-Code)」教學路徑:所有分析皆以 SPSS Modeler 的圖形化節點拖曳完成,不撰寫任何程式碼。演算法原理以「概念圖解+運動情境類比」方式說明,數學推導僅作理解輔助,不列入評量。每次上課採「課堂講授觀念及學生實機操作」。
二、針對非資訊背景學生之教學設計
零程式路徑 全程使用 SPSS Modeler 圖形化介面,以滑鼠拖曳節點、連線構成分析流程,不需撰寫 Python、R 或 SQL 程式碼。
三段式上機 每週上機分為三段:①單元負責學生示範→ ② 全班學生同步操作→ ③ 自行變化參數並觀察差異。
先看結果再懂原理 每章提供已建好的範例串流檔,學生先執行看到輸出,再回頭理解每個節點的作用,降低初學挫折。
運動語言解釋演算法 以體育運動情境類比抽象概念(如以「分組淘汰賽」理解決策樹、以「選手分型」理解集群),避免純數學陳述。
分組互助 分組教學,組內能力互補、上機時互相支援;文獻導讀以小組形式進行,降低個人負擔。
評量重點 評量重心在於「能否正確選用方法、正確解讀結果、正確轉譯為運動實務建議」,而非工具操作的熟練度或數學推導。
學習支援 提供圖解版節點速查表、每章操作步驟截圖手冊,以及課後線上諮詢時段。

三、學習目標
1. 能以運動專業語言說明資料探勘、機器學習與人工智慧之關係,並辨識體育運動場域中適用之問題類型。
2. 能依 CRISP-DM 流程獨立規劃並執行一項完整的體育運動資料探勘專案。
3. 能以 SPSS Modeler 圖形化介面建置資料流,完成資料匯入、清理、衍生、建模與評估(不需程式設計)。
4. 能說明各主要演算法(C5.0、C&R Tree、PCA、NN、BN、SVM、關聯規則、次序分析、K-means、Kohonen)的適用時機、參數意義與限制,並為自己的研究問題做出有依據的方法選擇。
5. 能以混淆矩陣、增益圖、ROC/AUC、輪廓係數等指標評估與比較模型,並辨識過度配適與資料洩漏。
6. 能批判性評析國際期刊之運動資料探勘研究,指出其方法論優劣並提出改進設計。
7. 能將分析結果轉譯為教練、選手、運動管理者可理解之決策建議,並落實資料倫理與個資保護。

四、教科書與參考書目
【主要教材】
《資料探勘:人工智慧與機器學習發展-以 SPSS Modeler 為範例》(書號 MP31904)。全書 16 章,本課程完整涵蓋。上機採「先做書上的例子、再做運動的例子」雙軌設計:第 3–14 章均使用隨書所附之範例資料與已建好的 .str 串流檔先行演練,再將同一套方法套用於體育運動資料集。
【輔助教材(教師自編)】
《體育運動資料探勘專案分析講義》CH01–CH16:每章含觀念直覺說明、SPSS Modeler 圖解操作步驟、體育運動應用實例、上機實作與作業,另附運動情境資料集與範例串流檔。
【參考書目】
‧ Han, J., Kamber, M., & Pei, J. Data Mining: Concepts and Techniques (3rd ed.). Morgan Kaufmann.
‧ Tan, P.-N., Steinbach, M., & Kumar, V. Introduction to Data Mining (2nd ed.). Pearson.
‧ Miller, T. W. Sports Analytics and Data Science. Pearson FT Press.
‧ Passfield, L., & Hopker, J. G. A mine of information: Can sports analytics provide wisdom from your data? International Journal of Sports Physiology and Performance.
‧ Van Eetvelde, H., et al. Machine learning methods in sport injury prediction and prevention: A systematic review. Journal of Experimental Orthopaedics.
(請尊重智慧財產權,不得非法影印教師指定之教科書籍)
五、教學要點概述
教材編選 ■自編教材 ■教科書作者提供 ■國際期刊論文選讀 ■運動情境範例資料集與串流檔
教學方法 ■投影片講述 ■板書講述 ■電腦教室實機操作(占課堂 50%) ■專題討論(Seminar) ■小組文獻導讀 ■個別專題指導
教學資源 ■課程網站 ■教材電子檔供下載 ■SPSS Modeler 實習環境 ■圖解操作手冊
■教科書隨書範例檔(MP31904_Example,Ch03–Ch14 各章附可直接執行之 .str 串流檔)
■教科書隨書投影片(MP31904_PPT,第 01–16 章) ■教師自編運動情境資料集
教學相關配合事項 ① 上課於電腦教室進行,學生亦可自備筆電並安裝 IBM SPSS Modeler(校園授權或試用版)。② 本課程不要求程式設計能力。③ 專題所使用之運動資料若涉及可識別個人資訊,須完成去識別化並說明資料取得之合法性;涉及人體研究者應通過研究倫理審查。

六、評量方法(博士班專案導向配分)
評量項目 配分 說明 對應目標
出席與課堂參與 10% 出席、課堂提問與討論、上機互助表現 1、6
上機實作作業 25% 每週課堂完成之 SPSS Modeler 串流與結果解讀,共 12 次,擇優 10 次計分 3、4、5
小組文獻導讀 15% 5 組×2 次;每組導讀 1 篇運動資料探勘國際期刊論文,評析其方法與限制 6
期中專題提案 20% 第 9 週書面提案(5 頁內)+ 口頭報告 8 分鐘:研究問題、資料來源、變項定義、分析計畫 1、2、7
期末專題分析報告 20% 完整書面報告,依期刊格式撰寫(緒論、方法、結果、討論、限制與應用建議) 2、5、7
期末口頭報告 10% 第 17–18 週個人公開報告 10 分鐘+問答 5 分鐘 6、7
※ 本課程不設期中考與期末考;以專題成果與實作表現取代紙筆測驗。演算法之數學推導不列入評量。
七、課程進度(115學年第1學期;每週 100 分鐘=講授 50 分鐘+上機 50 分鐘;日期依本校行事曆調整)
週次/日期 教科書
章節 【前 50 分鐘】講授重點 【後 50 分鐘】上機實作/運動應用實例
W1
260916 課程
導論 課程說明與評量方式;資料探勘能為體育運動做什麼;CRISP-DM 六階段 SPSS Modeler 安裝、介面導覽、執行第一個範例串流
★里程碑:分組、題目發想
W2
260923 CH1 資料探勘、機器學習與 AI 的關係;KDD 流程;探勘與統計檢定的差異 節點面板操作;建立「Excel→類型→表格→資料稽核」串流
★里程碑:確定專題方向
W3
260930 CH2 六大功能:分類、估計、預測、關聯、集群、描述;監督式與非監督式 欄位量測類型與角色設定;八個運動研究問題的功能判定演練
W4
261007 CH3 資料庫、資料倉儲與大數據 5V;運動資料的星狀綱要設計 多來源匯入與合併(Merge/Append/Aggregate):賽會成績+選手基本資料+體適能檢測三表整合
W5
261014 CH4 資料品質、遺漏值與離群值;標準化、分箱、衍生變項;抽樣與資料分割 青少年選手檢測資料清理;建立 ACWR 衍生變項;Partition 與 Balance
★里程碑:資料取得與清理完成
W6
261021 CH5 決策樹的直覺(像分組淘汰賽);熵與資訊增益;修剪與過度配適 C5.0 建模:由體能檢測預測選手競技層級;混淆矩陣與規則萃取
W7
261028 CH6 C&R Tree 與 C5.0 的差異;迴歸樹;代理分割;成本矩陣 運動傷害風險三級分類;訓練負荷對成績的迴歸樹;兩模型比較
W8
261104 CH7 為什麼要降維;主成分與因素分析;特徵值、陡坡圖、轉軸與命名 12 項體適能檢測的構面萃取;建立體能雷達圖與綜合指標
W9
261111 期中 期中專題提案口頭報告(每人 8 分鐘)與全班互評 教師個別回饋與資料/方法診斷
★期中考核:專題提案(20%)
W10
261118 CH8 類神經網路的直覺;隱藏層與激活函數;過度配適與早停;黑箱與可解釋性 由訓練量與生理指標預測 3000 公尺成績;過度配適曲線實作
W11
261125 CH9 機率思維與貝氏定理;貝氏網路結構;TAN 與 Markov Blanket 運動傷害成因網路建構與情境機率推論
W12
261202 CH10 支援向量機的直覺(找最寬的分隔帶);核函數;高維小樣本的優勢 由穿戴感測器特徵辨識羽球擊球動作類型
W13
261209 CH11 關聯規則:支持度、信賴度、增益;Apriori 與 Carma 排球得分戰術組合分析;運動中心課程與商品組合分析
W14
261216 CH12 次序分析:事件序列、時間窗、序列規則 籃球比賽事件序列分析;傷害前兆序列探勘
★里程碑:初步建模結果
W15
261223 CH13
CH14 K-means 集群與 k 值選擇;Kohonen 自組織映射與拓樸圖 選手型態分群與剖繪;運動參與者市場區隔;SOM 二維地圖視覺化
W16
261230 CH15
CH16 資料探勘與 AI/機器學習的發展趨勢;模型佈署、監控與再訓練;研究倫理 Auto Classifier/Auto Cluster 自動化建模與多模型比較;分析結果的教練端轉譯
W17
270106 期末
報告I 專題口頭報告與答辯(第 1–10 位,每人 10 分鐘+問答 5 分鐘) 同儕互評與教師講評
★期末考核:口頭報告
W18
270113 期末
報告II 專題口頭報告與答辯(第 11–20 位);專題論文化撰寫指導 書面報告繳交
★期末考核:書面報告(20%)

八、講授內容架構
PART I 資料探勘基礎與運動數據生態(CH1–CH2;W1–W3)
 ‧ 資料探勘、機器學習與人工智慧之概念界定與 KDD 流程
 ‧ CRISP-DM 六階段及其在體育運動研究之對應
 ‧ 運動數據來源:賽事統計、穿戴裝置、GPS/慣性感測、視訊追蹤、體適能檢測、會員與票務系統
 ‧ 六大資料探勘功能與運動問題之對應判準

PART II 運動資料工程與資料準備(CH3–CH4;W4–W5)
 ‧ 資料倉儲、大數據 5V 與運動資料的星狀綱要設計
 ‧ SPSS Modeler 資料流建置:來源、記錄選項、欄位選項節點
 ‧ 資料品質診斷:遺漏值、離群值、量尺一致性、跨年度量測漂移
 ‧ 衍生變項工程:訓練負荷(ACWR)、相對指標、變化率、左右不對稱指數
 ‧ 抽樣、資料分割與類別不平衡處理(先分割後平衡的正確順序)

PART III 監督式學習與運動預測建模(CH5–CH10;W6–W12)
 ‧ 決策樹 C5.0 與 C&R Tree:可解釋規則在選才與傷害分級之應用
 ‧ 主成分/因素分析:體適能構面萃取與降維
 ‧ 類神經網路:運動表現預測與非線性關係建模
 ‧ 貝氏網路:傷害成因之機率推論與情境模擬
 ‧ 支援向量機:動作辨識與高維小樣本之技術分類
 ‧ 模型評估:混淆矩陣、精確率/召回率、ROC-AUC、增益圖與過度配適防範

PART IV 非監督式學習與型態發現(CH11–CH14;W13–W15)
 ‧ 關聯規則:戰術組合、訓練處方組合與消費行為分析
 ‧ 次序分析:比賽事件序列、傷害前兆序列與訓練歷程
 ‧ K-means 集群:運動員型態分群與參與者市場區隔
 ‧ Kohonen/SOM:高維運動資料之視覺化拓樸分析

PART V AI/ML 發展趨勢與專案佈署(CH15–CH16;W16–W18)
 ‧ 整體學習與自動化建模(Auto Classifier/Auto Cluster)
 ‧ 模型佈署、效能監控、概念漂移與再訓練機制
 ‧ 分析結果之決策轉譯與教練端溝通
 ‧ 運動資料倫理、個資去識別化與研究倫理審查
 ‧ 專題成果之論文化:期刊格式撰寫與投稿策略

九、核心能力
1. 體育運動資料管理與資料準備能力
 整合賽事統計、穿戴裝置、體適能檢測與管理系統等多來源資料,以圖形化工具建立可重複執行的資料處理流程,診斷並處理資料品質問題,建立符合研究倫理與個資規範之分析資料集。
2. 資料探勘方法之正確選用與辯護能力
 能說明各演算法的適用條件、參數意義與限制,並依研究問題性質、資料型態與樣本規模,做出具學理依據的方法選擇並為其辯護。
3. SPSS Modeler 專案級實作能力(零程式)
 能獨立以圖形化節點建置模組化、可重現的分析串流,完成資料準備、建模、參數調校、模型比較與結果輸出,並管理專案版本與分析紀錄。
4. 模型評估與科學論證能力
 能運用適當評估指標與驗證設計檢驗模型效能,辨識過度配適、資料洩漏與抽樣偏誤,並以嚴謹論證區分統計關聯與因果推論。
5. 運動科學知識轉譯與決策支援能力
 能將模型結果轉化為教練、防護員、選手與運動管理者可執行之建議,並評估其實務可行性與潛在風險。
6. 學術研究與專案領導能力
 能批判性評析國際期刊之運動資料探勘文獻,獨立完成具投稿水準之專題研究報告,並具備規劃與領導產學合作分析專案之能力。
National Chung Cheng University Course Syllabus
Academic Year 2026, Semester 1
College of Education, Graduate Institute of Education, Doctoral Program — Physical Education and Sport Track
Course Title (Chinese) 資料探勘與大數據分析專題研究
Course Title (English) Seminar on Data Mining and Big Data Analytics
Instructor Lin Chin-Jung, Ph.D. / Professor
Credits / Hours 2 credits
Instructional Format Lecture and hands-on computer practice
Class Schedule Tuesday 10:10 AM–12:00 PM (actual schedule subject to the department's official timetable)
Classroom Computer lab (one workstation per student; IBM SPSS Modeler must be installed)
Prerequisites / Prior Preparation None. This course requires no programming or IT background; all operations are performed through a graphical user interface.
Course Attribute Doctoral seminar — conducted primarily as "Project-based Analytics"
I. Course Overview
This is a doctoral-level seminar course that uses Data Mining: The Development of Artificial Intelligence and Machine Learning — Illustrated with SPSS Modeler as its primary textbook, with all applied examples drawn from the domain of sport and exercise data analytics. Unlike the tool-operation orientation typical of master’s-level courses, this course is organized around “Project-based Analytics”: starting in Week 1, each student selects a research-worthy topic in sport data mining and advances it week by week through the six stages of CRISP-DM (the Cross-Industry Standard Process for Data Mining) — Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, and Deployment — culminating in a publication-quality project analysis report by the end of the semester.
Given that enrolled students come from a sport-science background and mostly have no prior programming or IT training, the course adopts a “No-Code” instructional path: all analyses are completed through SPSS Modeler’s graphical drag-and-drop nodes, without writing any code. Algorithmic principles are explained through “conceptual diagrams plus sport-context analogies”; mathematical derivations serve only as an aid to understanding and are not included in assessment. Each class session combines conceptual lecture with hands-on computer practice.
II. Instructional Design for Students without an IT Background
No-Code Path The entire course uses SPSS Modeler's graphical interface — analysis workflows are built by dragging and connecting nodes with the mouse; no Python, R, or SQL code is required.
Three-Stage Lab Session Each week's lab session is divided into three stages: (1) the student(s) assigned to that unit demonstrate → (2) the whole class follows along together → (3) students experiment with parameter variations and observe the differences.
See the Result First, Understand the Principle Second A pre-built example stream file is provided for every chapter; students first run it to see the output, then work backward to understand the function of each node — reducing the frustration often experienced by beginners.
Explaining Algorithms in the Language of Sport Abstract concepts are explained through sport-context analogies (e.g., using a single-elimination tournament bracket to understand decision trees, or athlete-type profiling to understand clustering), avoiding purely mathematical exposition.
Group-Based Mutual Support Instruction is conducted in groups, with complementary abilities within each group supporting one another during lab sessions; literature reviews are also conducted in groups to reduce the burden on individual students.
Assessment Focus Assessment focuses on whether students can correctly select a method, correctly interpret results, and correctly translate findings into practical sport recommendations — rather than on tool-operation fluency or mathematical derivation.
Learning Support Support includes an illustrated node quick-reference guide, a step-by-step screenshot manual for every chapter, and after-class online consultation hours.
III. Learning Objectives
1. Explain, in the language of sport science, the relationship among data mining, machine learning, and artificial intelligence, and identify the types of problems in the sport domain to which they apply.
2. Independently plan and execute a complete sport data mining project following the CRISP-DM process.
3. Build data streams using the SPSS Modeler graphical interface to complete data import, cleaning, derivation, modeling, and evaluation (no programming required).
4. Explain the appropriate use cases, parameter meanings, and limitations of the major algorithms (C5.0, C&R Tree, PCA, Neural Networks, Bayesian Networks, SVM, Association Rules, Sequence Analysis, K-means, and Kohonen), and make a well-justified method selection for one’s own research question.
5. Evaluate and compare models using metrics such as the confusion matrix, gains chart, ROC/AUC, and silhouette coefficient, and identify overfitting and data leakage.
6. Critically appraise international peer-reviewed studies in sport data mining, identify the strengths and weaknesses of their methodology, and propose improved designs.
7. Translate analytical results into decision recommendations that coaches, athletes, and sport managers can understand, while upholding data ethics and personal data protection.
IV. Textbook and References
Primary Textbook
Data Mining: The Development of Artificial Intelligence and Machine Learning — Illustrated with SPSS Modeler (Book No. MP31904). The book’s 16 chapters are covered in full. Lab sessions follow a dual-track design of “textbook example first, sport example second”: for Chapters 3–14, students first practice with the example data and pre-built .str stream files included with the textbook, then apply the same method to a sport dataset.
Supplementary Material (Instructor-Prepared)
Project Analytics in Sport Data Mining — Course Notes, Ch. 1–16: each chapter includes an intuitive conceptual explanation, illustrated SPSS Modeler operating steps, sport application examples, lab exercises and assignments, plus accompanying sport-context datasets and example stream files.
References
• Han, J., Kamber, M., & Pei, J. Data Mining: Concepts and Techniques (3rd ed.). Morgan Kaufmann.
• Tan, P.-N., Steinbach, M., & Kumar, V. Introduction to Data Mining (2nd ed.). Pearson.
• Miller, T. W. Sports Analytics and Data Science. Pearson FT Press.
• Passfield, L., & Hopker, J. G. A mine of information: Can sports analytics provide wisdom from your data? International Journal of Sports Physiology and Performance.
• Van Eetvelde, H., et al. Machine learning methods in sport injury prediction and prevention: A systematic review. Journal of Experimental Orthopaedics.
(Please respect intellectual property rights; illegal photocopying of the instructor-assigned textbook is prohibited.)
V. Teaching Overview
Course Materials ☑ Instructor-prepared materials ☑ Materials provided by the textbook author ☑ Selected international journal articles ☑ Sport-context example datasets and stream files
Teaching Methods ☑ Slide-based lecture ☑ Whiteboard lecture ☑ Hands-on practice in the computer lab (50% of class time) ☑ Seminar discussion ☑ Group literature review presentations ☑ Individual project mentoring
Teaching Resources ☑ Course website ☑ Downloadable electronic course materials ☑ SPSS Modeler practice environment ☑ Illustrated operating manual
☑ Textbook example files (MP31904_Example — ready-to-run .str stream files for Ch. 3–14)
☑ Textbook slides (MP31904_PPT — Ch. 1–16) ☑ Instructor-prepared sport-context datasets
Course Logistics ① Classes are held in the computer lab; students may also bring their own laptop with IBM SPSS Modeler installed (campus license or trial version). ② No programming ability is required for this course. ③ If the sport data used in a project contains personally identifiable information, it must be de-identified, and the legality of how the data was obtained must be explained; projects involving human subjects research must pass research ethics review.
VI. Assessment Methods (Doctoral, Project-Oriented Weighting)
Assessment Item Weight Description Objective(s)
Attendance & Class Participation 10% Attendance, in-class questions and discussion, mutual support during lab sessions 1, 6
Weekly Lab Assignments 25% SPSS Modeler streams and result interpretation completed in class each week; 12 sessions total, best 10 counted 3, 4, 5
Group Literature Review Presentations 15% 5 groups × 2 sessions; each group presents one international peer-reviewed sport data mining article, critiquing its methods and limitations 6
Midterm Project Proposal 20% Week 9 written proposal (max. 5 pages) + 8-minute oral presentation: research question, data source, variable definitions, and analysis plan 1, 2, 7
Final Project Analysis Report 20% Complete written report in journal-article format (Introduction, Methods, Results, Discussion, Limitations, and Practical Recommendations) 2, 5, 7
Final Oral Presentation 10% Individual public presentation in Weeks 17–18, 10 minutes + 5-minute Q&A 6, 7
Note: This course has no midterm or final written examination; project outcomes and hands-on performance replace paper-and-pencil testing. Mathematical derivations of algorithms are not included in the assessment.
VII. Course Schedule (Academic Year 2026, Semester 1; 100 minutes/week = 50 min lecture + 50 min lab; dates subject to the university calendar)
Week / Date Textbook Ch. First 50 min — Lecture Focus Last 50 min — Lab / Sport Application
W1
2026/09/16 Course
Introduction Course overview and grading policy; what data mining can do for sport; the six stages of CRISP-DM SPSS Modeler installation, interface orientation, running the first example stream
★ Milestone: Group formation, topic brainstorming
W2
2026/09/23 Ch.1 The relationship among data mining, machine learning, and AI; the KDD process; how mining differs from statistical hypothesis testing Node-panel operations; building an “Excel → Type → Table → Data Audit” stream
★ Milestone: Project direction finalized
W3
2026/09/30 Ch.2 The six core functions: classification, estimation, prediction, association, clustering, description; supervised vs. unsupervised learning Field measurement level and role settings; hands-on determination of the appropriate function for eight sport research questions
W4
2026/10/07 Ch.3 Databases, data warehouses, and the 5 V’s of big data; star-schema design for sport data Multi-source import and merging (Merge/Append/Aggregate): integrating competition results, athlete profile data, and fitness test data across three tables
W5
2026/10/14 Ch.4 Data quality, missing values and outliers; standardization, binning, derived variables; sampling and data partitioning Cleaning youth-athlete testing data; creating the ACWR derived variable; Partition and Balance nodes
★ Milestone: Data acquisition and cleaning completed
W6
2026/10/21 Ch.5 The intuition behind decision trees (like a single-elimination bracket); entropy and information gain; pruning and overfitting C5.0 modeling: predicting athletes’ competitive level from fitness test results; confusion matrix and rule extraction
W7
2026/10/28 Ch.6 Differences between C&R Tree and C5.0; regression trees; surrogate splits; cost matrices Three-level classification of sport injury risk; regression tree of training load on performance; comparing the two models
W8
2026/11/04 Ch.7 Why reduce dimensionality; principal components and factor analysis; eigenvalues, scree plots, rotation and naming Extracting dimensions from 12 fitness test items; building a fitness radar chart and composite index
W9
2026/11/11 Midterm Oral presentation of midterm project proposals (8 minutes per student) with whole-class peer feedback Individual instructor feedback and data/method diagnosis
★ Midterm assessment: Project proposal (20%)
W10
2026/11/18 Ch.8 The intuition behind neural networks; hidden layers and activation functions; overfitting and early stopping; the black-box problem and interpretability Predicting 3,000 m performance from training volume and physiological indicators; hands-on demonstration of overfitting curves
W11
2026/11/25 Ch.9 Probabilistic thinking and Bayes’ theorem; Bayesian network structure; TAN and the Markov blanket Building a sport-injury causation network and scenario-based probabilistic inference
W12
2026/12/02 Ch.10 The intuition behind support vector machines (finding the widest separating margin); kernel functions; advantages for high-dimensional, small-sample data Identifying badminton stroke types from wearable-sensor features
W13
2026/12/09 Ch.11 Association rules: support, confidence, lift; Apriori and CARMA Volleyball scoring-tactic combination analysis; market-basket analysis of sport-center program and product bundling
W14
2026/12/16 Ch.12 Sequence analysis: event sequences, time windows, sequence rules Basketball game event-sequence analysis; mining injury-precursor sequences
★ Milestone: Preliminary modeling results
W15
2026/12/23 Ch.13
Ch.14 K-means clustering and selecting k; Kohonen self-organizing maps and topological mapping Athlete-type clustering and profiling; market segmentation of sport participants; 2-D visualization with SOM
W16
2026/12/30 Ch.15
Ch.16 Trends in the development of data mining and AI/machine learning; model deployment, monitoring, and retraining; research ethics Automated modeling and multi-model comparison with Auto Classifier / Auto Cluster; translating analytical results for coaches
W17
2027/01/06 Final
Presentations I Oral project presentations and defense (students 1–10, 10 minutes each + 5-minute Q&A) Peer feedback and instructor commentary
★ Final assessment: Oral presentation
W18
2027/01/13 Final
Presentations II Oral project presentations and defense (students 11–20); guidance on preparing the project for journal publication Submission of the written report
★ Final assessment: Written report (20%)
VIII. Lecture Content Structure
PART I — Foundations of Data Mining and the Sport Data Ecosystem (Ch. 1–2; Weeks 1–3)
• Defining data mining, machine learning, and artificial intelligence; the KDD process
• The six stages of CRISP-DM and their correspondence to sport research
• Sources of sport data: competition statistics, wearable devices, GPS/inertial sensors, video tracking, fitness testing, membership and ticketing systems
• The six core data mining functions and criteria for matching them to sport questions
PART II — Sport Data Engineering and Data Preparation (Ch. 3–4; Weeks 4–5)
• Data warehousing, the 5 V’s of big data, and star-schema design for sport data
• Building SPSS Modeler data streams: Source, Record Options, and Field Options nodes
• Data quality diagnostics: missing values, outliers, scale consistency, cross-year measurement drift
• Derived-variable engineering: training load (ACWR), relative indicators, rate of change, left–right asymmetry indices
• Sampling, data partitioning, and handling class imbalance (the correct order: partition first, then balance)
PART III — Supervised Learning and Sport Predictive Modeling (Ch. 5–10; Weeks 6–12)
• Decision trees — C5.0 and C&R Tree: interpretable rules for talent identification and injury grading
• Principal component / factor analysis: extracting fitness dimensions and dimensionality reduction
• Neural networks: predicting sport performance and modeling nonlinear relationships
• Bayesian networks: probabilistic inference on injury causation
• Support vector machines: motion recognition and classification with high-dimensional, small-sample data
• Model evaluation: confusion matrix, precision/recall, ROC-AUC, gains charts, and guarding against overfitting
PART IV — Unsupervised Learning and Pattern Discovery (Ch. 11–14; Weeks 13–15)
• Association rules: tactic combinations, training-prescription bundling, and consumer behavior analysis
• Sequence analysis: game event sequences, injury-precursor sequences, and training histories
• K-means clustering: athlete-type segmentation and participant market segmentation
• Kohonen/SOM: topological visualization of high-dimensional sport data
PART V — AI/ML Trends and Project Deployment (Ch. 15–16; Weeks 16–18)
• Ensemble learning and automated modeling (Auto Classifier / Auto Cluster)
• Model deployment, performance monitoring, concept drift, and retraining mechanisms
• Translating analytical results into decisions and communicating with coaches
• Sport data ethics, de-identification of personal data, and research ethics review
• Turning project outcomes into a publishable paper: journal formatting and submission strategy
IX. Core Competencies
1. Sport Data Management and Preparation. Integrate multi-source data — competition statistics, wearable devices, fitness testing, and management systems — using graphical tools to build reproducible data-processing workflows; diagnose and resolve data quality issues; and construct analysis-ready datasets that comply with research ethics and personal data protection requirements.
2. Sound Selection and Justification of Data Mining Methods. Explain the applicable conditions, parameter meanings, and limitations of each algorithm, and make — and defend — a theoretically grounded choice of method based on the nature of the research question, data type, and sample size.
3. Project-Level SPSS Modeler Implementation (No-Code). Independently build modular, reproducible analysis streams using graphical nodes to complete data preparation, modeling, parameter tuning, model comparison, and output generation, while managing project versions and analysis records.
4. Model Evaluation and Scientific Reasoning. Use appropriate evaluation metrics and validation designs to test model performance, identify overfitting, data leakage, and sampling bias, and rigorously distinguish statistical association from causal inference.
5. Translation of Sport Science Knowledge and Decision Support. Translate model results into actionable recommendations for coaches, athletic trainers, athletes, and sport managers, while assessing their practical feasibility and potential risks.
6. Academic Research and Project Leadership. Critically appraise the international sport data mining literature, independently complete a publication-quality project research report, and demonstrate the ability to plan and lead industry–academia collaborative analytics projects.

二、課程大綱說明文件260830資料探勘與大數據分析專題研究授課大綱.docx
260901CCU_Syllabus_Seminar on Data Mining and Big Data Analytics_EN.docx
三、教材編選
四、教學教法
五、評量工具
請尊重智慧財產權,不得非法影印教師指定之教科書籍