晶文摘人工智慧的應用實作
作者: 夏肇毅
初稿: 20220819
自然語言處理(Natural language processing)
http://cubicpower.idv.tw/cubicnotes/notes-0000038.html
文字探勘(Text Mining)
Google Tensorflow: Text
Sentiment analysis- IMDB large movie review dataset
Basic text classification
https://www.tensorflow.org/tutorials/keras/text_classification?hl=zh-tw
基本文本分類
情緒分析
下載並探索 IMDB 數據集
加載數據集
準備數據集進行訓練
配置數據集以提高性能
創建模型
損失函數和優化器
訓練模型
評估模型
創建隨時間變化的準確度和損失圖
導出模型
對新數據的推論
練習:關於 Stack Overflow 問題的多類分類
Word embeddings
https://www.tensorflow.org/text/guide/word_embeddings?hl=zh-tw
詞嵌入
將文本表示為數字
One-hot 編碼
用唯一的數字編碼每個單詞
詞嵌入
設置
下載 IMDb 數據集
使用嵌入層
文本預處理
創建分類模型
編譯和訓練模型
檢索經過訓練的詞嵌入並將它們保存到磁碟
可視化嵌入
Text classification with an RNN
https://www.tensorflow.org/text/tutorials/text_classification_rnn?hl=zh-tw
使用 RNN 進行文本分類
設置
設置輸入管道
創建文本編碼器
創建模型
訓練模型
堆疊兩個或多個 LSTM 層
Classify text with BERT
https://www.tensorflow.org/text/tutorials/classify_text_with_bert?hl=zh-tw
使用 BERT 對文本進行分類
關於 BERT
情緒分析
從 TensorFlow Hub 加載模型
選擇一個 BERT 模型進行微調
預處理模型
使用 BERT 模型
定義你的模型
模型訓練
損失函數
優化器
加載 BERT 模型並進行訓練
評估模型
繪製隨時間變化的準確性和損失
導出推理
Search: LogisticRegression
[Day 9] 邏輯迴歸(Logistic Regression) - iT 邦幫忙
https://ithelp.ithome.com.tw › articles
[Python實作]邏輯斯迴歸模型Logistic Regression - PyInvest
https://pyecontech.com › 2020/02/06 › python_logistic...
Search: kaggle 登入
https://nancysw.medium.com › 兩種取用-kaggle-資料...
Search: 聊天機器人 python
在Python 中實作對話型聊天機器人 - WANcatServer
https://wancat.cc › post › python-chatbot-context
Search: chatbot python github
zake7749/Chatbot: 基於向量匹配的情境式聊天機器人 - GitHub
https://github.com › zake7749 › Chatbot
Mianbot
匹配示例
環境需求
使用方式
聊天機器人
計算匹配度
規則格式
問答測試用資料集
Building a Simple Chatbot from Scratch in Python (using NLTK)
https://github.com › parulnith › Building-...
Building a Simple Chatbot from Scratch in Python (using NLTK)
NLP
Import necessary libraries
Downloading and installing NLTK
Installing NLTK Packages
Reading in the corpus
Tokenisation
Preprocessing
Keyword matching
Generating Response
Bag of Words
TF-IDF Approach
Cosine Similarity
Day 13:快速完成一個『對話機器人』(ChatBot) -- 續
Search: 反洗錢 模型 python
https://www.ithome.com.tw › news
(Top 1% Solution)玉山人工智慧公開挑戰賽2020夏季賽 - Medium
https://medium.com › 玉山人工智慧公開挑戰賽2020夏...
Search: 防詐欺模型 python
靠AI扮演壞人來練兵!SAS揭露用GAN設計金融防詐欺模型的新 ...
https://www.ithome.com.tw › news
步驟一、構建Amazon Fraud Detector 模型
https://pages.awscloud.com › Tech-blog_Amazon-Frau...
Search: 詐欺 python
https://www.796t.com › content
資料準備: 來源於Kaggle
準備並初步檢視資料集
時間序列下的交易發生頻率(分為詐騙和正常)
詐騙和正常交易交易金額的頻率分佈
各特徵和因變數的關係
用邏輯迴歸方法對信用卡資料進行建模分析
信用卡詐騙分析-不平衡資料分析與處理kernel翻譯-完整版
https://medium.com › 機器學習知識歷程 › 信用卡詐騙...
預處理
縮放和分配 Scaling and Distributing
拆分數據 Splitting the Data(從原始DataFrame)
隨機欠採樣和過採樣
分佈和相關性 Distributing and Correlating
異常檢測 Anomaly Detection
降維和分群 Dimensionality Reduction and Clustering (t-SNE)
分類器 Classifiers
更深入地了解邏輯回歸 A Deeper Look into Logistic Regression
使用SMOTE進行過採樣 Oversampling with SMOTE
測試
使用邏輯回歸進行測試 Test Data with Logistic Regression
神經網絡測試(欠採樣與過採樣)Neural Networks Testing (Undersampling vs Oversampling)
Part 5. Imbalanced Data 不平衡資料 - iT 邦幫忙
https://ithelp.ithome.com.tw › articles
評估指標
A. Confusion Matrix 混淆矩陣
B. Precision and Recall 精確率與召回率
C. F1 score
D. ROC(Receiver Operating Characteristic) 接收者操作特徵曲線
曲線下面積稱 Area Under Curve (AUC)
重組資料
A. Oversampling 過採樣
A1. SMOTE (Synthetic Minority Oversampling Technique)
A2. Border Line SMOTE
B. Undersampling 欠採樣
B1. Tomek Link
B2. Edited Nearest Neighbor
注意事項
A. 先切分資料,再對訓練資料採樣。
B. 常透過交叉驗證控制過擬合。
C. 觀察少數樣本與多數樣本分布情形。
Search: Credit Fraud || Dealing with Imbalanced Datasets
Credit Fraud || Dealing with Imbalanced Datasets - Kaggle
https://www.kaggle.com › janiobachmann
Credit Fraud || Dealing with Imbalanced Datasets
https://www.kaggle.com/janiobachmann/credit-fraud-dealing-with-imbalanced-datasets
https://www.kaggle.com/code/janiobachmann/credit-fraud-dealing-with-imbalanced-datasets/notebook
Search: 信用評分 python
基於Python的信用評分模型開發-附資料和程式碼 - 古詩詞庫
https://www.gushiciku.cn › zh-tw
專案流程
資料獲取
資料預處理
缺失值處理
異常值處理
資料切分: 分成 訓練集和測試集
探索性分析
變數選擇
分箱處理
WOE
相關性分析和IV篩選
模型分析
WOE轉換
Logisic模型建立
模型檢驗
信用評分
自動評分系統
Search: GiveMeSomeCredit
jwu424/GiveMeSomeCredit - GitHub
https://github.com › jwu424 › GiveMeSo...
信用限額: line of credit,credit line.
WOE(Weight of Evidence)證據權重,常用於特徵變換,
IV(Information Value)資訊價值,或者資訊量,用來衡量特徵的預測能力。
1. WOE describes the relationship between a predictive variable and a binary target variable.
2. IV measures the strength of that relationship.
1. WOE 描述了預測變量和二元目標變量之間的關係。
2. IV 衡量這種關係的強度。
Search: woe iv
Search: woe iv python
https://arsene5240.medium.com › woe值及iv值-python...
Search: 精準行銷 python
機器學習系列五:誰會簽約?以「精準行銷模型」評估顧客帶來 ...
https://medium.com › marketingdatascience › 機器學習...
機器學習X 精準行銷KDD 2.0程序:【內部資料】實案應用(附 ...
https://medium.com › marketingdatascience › 機器學習...
Search: 客戶分群 python
Python用K-means聚類演算法進行客戶分群的實現 - 程式人生
https://www.796t.com › article
Day 02:客戶分群(Customer Segmentation) -- 那些客戶是VIP?
客戶分群(Customer Segmentation) -- 那些客戶是我的VIP? (續)
https://ithelp.ithome.com.tw › articles
Search: 文字關聯 python
【python資料探勘課程】二十四.KMeans文字聚類分析互動百科 ...
https://hackmd.io › python-bigdata-02
文字雲
結巴
分詞 斷詞
詞頻
停用詞
Search: Python 文字探勘
文件探勘(Text Mining) — 把文字用數字表示 - Medium
https://medium.com › 企鵝也懂程式設計 › 文件探勘-te...
Text Mining & 網路爬蟲web crawler | Google新聞與文章文字雲
https://jamleecute.web.app › 網路爬蟲-web-crawler-text...
Search: Python 文字探勘 應用
[Python機器學習]-自動判斷留言正負評(運用BERT model)with ...
https://medium.com › python機器學習-google我的商...
Search: Python 文字探勘 技術
https://dannypheobe.blogspot.com › 2016/07 › python
10 Text Mining - LearnPython - GitBook
https://datasciencetw.gitbook.io › python › 10-text-mini...
Google Tensorflow: Text
Sentiment analysis- IMDB large movie review dataset
Basic text classification
https://www.tensorflow.org/tutorials/keras/text_classification?hl=zh-tw
Word embeddings
https://www.tensorflow.org/text/guide/word_embeddings?hl=zh-tw
Text classification with an RNN
https://www.tensorflow.org/text/tutorials/text_classification_rnn?hl=zh-tw
Classify text with BERT
https://www.tensorflow.org/text/tutorials/classify_text_with_bert?hl=zh-tw