刚刚开始学习R语言,学校的课程不是很系统,仅作为学习笔记参考。刚接触R时个人感觉R很像SQL的查询语句和MatLab的plot功能的结合,有过SQL经验的同学应该很好上手,不过R的Document感觉没有MATLAB做的直观,查阅资料有一定困扰。根据学校课程和各路资料汇总,不定时更改与更新。
利用conda创建了一个新的环境,安装Jupyter Lab(或者Jupyter Notebook),再安装R Essentials。activate 环境后在JupyterLab中使用R Kernel。
点此网址下载道琼斯指数dataset
#load required library library(tidyverse) #prerequisite: download the provided dataset and put it in to "path_to_downloaded_dataset" directory #now, import the dataset stocks <- read.table("self_name/dow_jones_index.data", header = TRUE, sep = ",")这里首先需要下载文件(两种方法): 1.可直接从网页下载到self_name这个自定义文件夹(推荐)。 2.利用download.file(url, destfile):url为网页地址,destfile为保存地址,默认工作目录,file为保存文件名。
#get list of companies companies <- unique(stocks$stock, incomparables = FALSE)这里unique(x, incomparables = FALSE, …)来剔除重复项,和SQL类似。x可是是向量,dataset,NULL或者数组;incomparables指明不可被比较的向量,一般为FALSE,表示所有值都可被比较。
#number of companies ncompanies <- length(companies) #declare correlation matrix correlations <- matrix(data = NA, nrow = ncompanies, ncol = ncompanies, byrow = FALSE, dimnames = NULL)官方:matrix(data = NA, nrow = 1, ncol = 1, byrow = FALSE, dimnames = NULL) 其中:没有的数据用NA来填充;行列数各为1;byrow为FALSE说明用列来填充(If FALSE (the default) the matrix is filled by columns, otherwise the matrix is filled by rows.);dimnames:各个维度的名字(即行列名称)
#compute pair-wise PEARSON correlation coefficient for (i in 1:ncompanies){ icompany <- arrange(filter(select(stocks, stock, date, volume), stock == companies[i]), date) for (j in i:ncompanies){ jcompany <- arrange(filter(select(stocks, stock, date, volume), stock == companies[j]), date) correlations[i,j] <- cor(icompany$volume, jcompany$volume, method = "pearson") correlations[j,i] <- correlations[i,j] } }对每一支股票stock_i(company),利用cor计算与另一支stock_j股票的相关系数。cor(x, y = NULL, use = “everything”, method = c(“pearson”, “kendall”, “spearman”)) 注:
pearson:(default):主要用于判断两连续变量是线性还是非线性spearman:如果变量是排序(ranked)的或是非线性相关kendall:反应分类变量相关性