Showing posts with label Social Network Analysis. Show all posts
Showing posts with label Social Network Analysis. Show all posts

Monday, December 23, 2013

An Actuarial Application of Social Network Analysis

While the article from The Future Actuary linked to below doesn't really mention social network analysis (SNA), it speaks to insurance products designed to protect you from loss, personal liability, or related litigation issues involving social media. I'm sure there has already been plenty of work done in this area and if I get the chance I'll provide some updates. I would think knowledge of SNA would be critical i.e what is the relationship between centrality measures and network structure and risk? How can this help price products designed to protect individuals and businesses?

"As a social media entrepreneur, Matt Mullenweg put it: “we’re only one click away from a billion people.” While these technological advancements have made ‘getting connected’ one of the top priorities for employers and individuals, it doesn’t come without its perils. The increasing risks of legality/operation, security, and defamation are also just one click away. Maybe it’s time for social media risk management plans to take off. "

Tuesday, September 3, 2013

GMM, Endogeneity, SNA, Viral Marketing, and Causal Inference



 In the article Impact of social network structure on content propagation: A study using YouTube data” the authors investigate the relationship between socioemetric measures like degree centrality with diffusion of videos across a network.  In other words, they wanted to know if there was a causal relationship between network properties of those that share videos and the likelihood that a video would become viral.  What first interested me about this article was that it was a very good example of an application of social network analysis and viral seeding.  However, it also provides some very good examples of applications related to generalized method of moments,  instrumental variables, unobserved heterogeneity and endogeneity, and causal inference.  I previously was not aware of the GMM style of dynamic panel data models that instrument with lags, which is apparently quite popular in many econometric applications (see references below).

As the authors point out, any model that relates network properties to the outcome of video dissemination requires a careful estimation strategy if we are interested in making causal inferences. They identity several sources of endogeneity and unobserved heterogeneity.  If we are trying to infer dissemination based on one’s position in the network, we have to consider that other unobserved factors related to network position and video type could also impact dissemination.  It may be the case that all we are trying to do is predict video shares based on network position,  and perhaps that is OK as long as these correlations hold over time.

In contrast, if we want to make causal inferences, these types of endogeneity must be accounted for and also make econometric estimation difficult. In this case what we really want to estimate is the independent causal effect of network position on video shares, so we are interested only in the ‘quasi-experimental’ variation in network position.

A natural solution involves an instrumental variables approach, but the challenge of finding an ‘external’ instrument that is correlated with network and video properties of interest, but uncorrelated with unobserved effects is rather daunting. Ultimately the authors propose a generalized method of moments dynamic panel estimator using lagged variables as instruments.  

References:

 Anderson, T. W., & Hsaio, C. (1981). Estimation of dynamic models with error components. Journal of the American Statistical Association, 76(375), 598–606.

Arellano, M., & Bond, S. (1991). Some tests of specification for panel data: Monte Carlo evidence and an application to employment equations. The Review of Economic Studies, 58, 277–97.

DYNAMIC PANEL DATA MODELS:
A GUIDE TO MICRO DATA METHODS AND PRACTICE
Stephen Bond
THE INSTITUTE FOR FISCAL STUDIES
DEPARTMENT OF ECONOMICS, UCL
cemmap working paper CWP09/02

Impact of social network structure on content
propagation: A study using YouTube data
Quant Mark Econ (2012) 10:111150
Hema Yoganarasimhan

Friday, May 31, 2013

References: Social Network Analysis and Student Integration and Persistence


See also: An Introduction to SNA using R and NetDraw, SNA & Predictive Modeling, Using Twitter to Demonstrate Basic Concepts from SNA,  An Introduction to SNA with Applications

Social Psychology of Education
June 2012, Volume 15, Issue 2, pp 165-180
A social network analysis of student retention using archival data
James E. Eckles,
Eric G. Stradley 


This study attempts to determine if a relationship exists between first-to-second-year retention and social network variables for a cohort of first-year students at a small liberal arts college. The social network is reconstructed using not survey data as is most common, but rather using archival data from a student information system. Each student is given a retention score and an attrition score based on the behavior of their immediate relationships in the network. Those scores are then entered into a logistic regression that includes tradition background and performance variables that are traditionally significantly related to retention. Students' friends' retention and attrition behaviors are found to have a greater impact on retention that any background or performance variable.


JOURNAL OF COLLEGE STUDENT RETENTION, Vol. 4(1) 39-52, 2002-2003
THE ROLE OFSOCIAL SUPPORT NETWORK
IN COLLEGE PERSISTENCE AMONG
FRESHMAN STUDENTS
MICHAEL P. SKAHILL
University of San Francisco, California

 Used social network analysis to examine the role of social support networks in student persistence among residential and commuter students. Found that commuter students are less likely to persist, while residential students who reported making greater numbers of new friends with connections to the school also reported attaining personal and academic goals at a significantly greater rate.

The Journal of Higher Education   
Vol. 71, No. 5, Sep. - Oct., 2.
Ties That Bind: A Social Network Approach to Understanding Student Integration and Persistence
Scott L. Thomas

This study examined the social networks of college students and how such networks affect student commitment and persistence. The study's theoretical framework was based on application of the social network paradigm to Tinto's Student Integration Model, in which a student's initial commitment is modified over time as a result of the student's integration into the campus community. Freshmen enrolled for the spring 1993 semester responded (322 of 379) to the First-Year Experiences Survey, which involved identifying students with whom they frequently spoke and the dimensions on which they related to these students. Results were compared with enrollment data for the fall 1993 semester to identify students returning for their sophomore year. The largest effect on persistence was associated with the number of nominations received from other students, and this factor operated indirectly through enhanced social integration, institutional commitment, and intention. Overall, students with broader, well-connected networks were more likely to persist, whereas students with a higher proportion of ties falling within their social peer group were less likely to persist.

Physics Education Research Conference 2009
Part of the PER Conference series
Ann Arbor, Michigan: July 29-30, 2009
Volume 1179, Pages 105-108
Investigating Student Communities with Network Analysis of Interactions in a Physics Learning Center written by Eric Brewe, Laird H. Kramer, and George O'Brien 

 Developing a sense of community among students is one of the three pillars of an overall reform effort to increase participation in physics, and the sciences more broadly, at Florida International University. The emergence of a research and learning community, embedded within a course reform effort, has contributed to increased recruitment and retention of physics majors. Finn and Rock [1] link the academic and social integration of students to increased rates of retention. We utilize social network analysis to quantify interactions in Florida International University's Physics Learning Center (PLC) that support the development of academic and social integration,. The tools of social network analysis allow us to visualize and quantify student interactions, and characterize the roles of students within a social network. After providing a brief introduction to social network analysis, we use sequential multiple regression modeling to evaluate factors which contribute to participation in the learning community. Results of the sequential multiple regression indicate that the PLC learning community is an equitable environment as we find that gender and ethnicity are not significant predictors of participation in the PLC. We find that providing students space for collaboration provides a vital element in the formation of supportive learning community.

Monday, April 15, 2013

SNA & Learning Communities

This week I'm at the CPE Student success summit. Learning communities are an ongoing theme at the conference. This article is one of the few that I've found that uses social network analysis metrics to investigate student learning communities. Are centrality measures good indicators of integration? 

Phys. Rev. ST Physics Ed. Research 8, 010101 (2012) [9 pages]

Investigating student communities with network analysis of interactions in a physics learning center

ABSTRACT

"Developing a sense of community among students is one of the three pillars of an overall reform effort to increase participation in physics, and the sciences more broadly, at Florida International University. The emergence of a research and learning community, embedded within a course reform effort, has contributed to increased recruitment and retention of physics majors. We utilize social network analysis to quantify interactions in Florida International University's Physics Learning Center (PLC) that support the development of academic and social integration. The tools of social network analysis allow us to visualize and quantify student interactions and characterize the roles of students within a social network. After providing a brief introduction to social network analysis, we use sequential multiple regression modeling to evaluate factors that contribute to participation in the learning community. Results of the sequential multiple regression indicate that the PLC learning community is an equitable environment as we find that gender and ethnicity are not significant predictors of participation in the PLC. We find that providing students space for collaboration provides a vital element in the formation of a supportive learning community."

http://prst-per.aps.org/abstract/PRSTPER/v8/i1/e010101 

Tuesday, November 20, 2012

An Intuitive Approach to Eigenvector Centrality using SAS IML

A network can be thought of in terms of graph theory as a set of vertices connected by ties. The vertices can include individuals, teams, government agencies, organizations, facebook group members, topics, patents,etc. (Coulon,2005). The ties, represented as lines connecting the vertices are referred to as ‘edges.’ Edges therefore indicate the connections between individuals, organizations etc. Vertices and connections can be represented by what’s referred to as an adjacency matrix ‘A’, where Aij = 1 if there is an edge between vertices ‘i’ and ‘j’, and Aij = 0 otherwise. Adjacency matrices are symmetric, in that Aij= Aji. For example, the following adjacency matrix corresponds to the network depicted below:



Who’s who in a network?
The role a given vertice plays in the cohesiveness of the network (its ‘importance’ or ‘centrality’ )can be assessed by a number of social network metrics. 

Degree Centrality is simply the number of connections (or edges) a vertex has to other vertices. Its clear from counting that #1’s degree centrality is 3, while vertex #4 has a measure of degree centrality equal to 1. 

In terms of matrix algebra, if we can obtain this result as a vector (one entry per vertice) as follows:
x= (1,1,1,1)  degree = A’x 

Eigenvector Centrality is a measure that reflects the fact that not all connections are equal, and in fact, connections to people that are more influential are more important (Newman, 2012).  Mathematically, the measure of eigenvector centrality is derived from the weights that consist of the values from the leading eigenvector  (the eigenvector associated with the largest eigenvalue) of the adjacency matrix depicting the network. 

Recall, the value λ is an eigenvalue (or vector of eigenvalues) of the matrix A, and x is an eigenvector of A if the following equation holds:

λ x = A x

So, for the network depicted above, the centrality measures for each vertex = 1,2,3,4 can be derived from the components (x1,x2,x3,x4) that comprise the eigenvector x, and what we are interested in for this metric is the leading eigenvector (associated with the largest eigenvalue of A). 

Eigenvectors and eigenvalues can be derived by solving:

(A-λ I)x = 0
This metric can be found using canned functions in SAS and R. In SAS IML, the eigen function will allow you to obtain the leading eigenvector  x =( 0.6116285,0.5227207,0.5227207,0.2818452). 

A Path Centric Approach to Centrality

One way of looking at centrality in terms of connectivity is counting the number of paths emanating from a vertice to other vertices in the network.  For example, degree represents the number of paths of length 1 to other vertices in the network.  To look at this further, lets define the matrix B as:
B = A + I, where I is the identity matrix of A. The elements of B represent the number of paths of length 1 from a vertice ‘i’ to another vertice ‘j’. The row sums of B tell us the total number of paths of length one emanating for a given vertice (counting each vertice itself as 1). We can store this sum for each vertice in a vector ‘d’ obtained by multiplyting B’x



This gives us B’x = d1 = (4,3,3,2)

Now let’s define the matrix power of such that B2 = B*B


Each entry in B2 represents the number of paths of length 2 from vertice i to vertice j.
The sum  d2 = B2 *x = d2 = (12,10,10,6) is the total number of paths of length 2 originating from each corresponding vertice.

So Bk is a matrix with elements indicating the number of paths of length k between vertices i and j. The product Bk*x gives the sum of all paths of length k emanating from a given vertice. What we find is that as we consider longer and longer paths (as k becomes large) the normalized vector:

v = Bk*x / || Bk*x||   will approach the leading eigenvector of B (which consequently is also the leading eigenvector of A)

v1 = B1*x / || Bk*x||    = (0.6488857,0.4866643,0.4866643,0.3244428)
v5 = B5*x / || Bk*x||   =  (0.6121647,0.521951,    0.521951,0.2835289)
v10 = B10*x / || Bk*x||   =( 0.6116348,0.5227115,0.5227115,0.2818657)

Recall the leading eigenvector of A:

 x = ( 0.6116285,0.5227207,0.5227207,0.2818452)

So, what we have is a notion of eigenvector centrality as a metric that captures the number of paths emanating from a given vertice. A vertice that is connected to other vertices that also have lots of connections will have a greater number of k-length paths originating from it. This is information is reflected by eigenvector centrality. It turns out that the iterative  calculation of v above is essentially  the power iteration algorithm for calculating the leading eigenvector of a matrix B.

The following code from SAS IML and output provides an illustration of the preceding discussion.  In addition I have provided the function written by Rick Wicklin that implements the power iteration algorithm used to obtain the leading eigenvector and eigenvalue of a given matrix A.

REFERENCES:
Justification and Application of Eigenvector Centrality by Leo Spizzirri     https://www.math.washington.edu/~morrow/336_11/papers/leo.pdf

The power method: compute only the largest eigenvalue of a matrix.  The Do Loop. By  Rick Wicklin.

SAS IML MATRIX PROGRAMMING: 
 
proc iml;

*specify adjacency matrix A;
A = {0 1 1 1,
     1 0 1 0,
      1 1 0 0,
      1 0 0 0};

x = {1,1,1,1};

* Degree centraltiy;

d1 = A*x;
print(A);
print(x);
print(d1);

call eigen(u,v,A);
print v;

B = A + I(4);
print B;

d1 = B*x;
print d1;

B2 = B**2;
print B2;

* Degree Centrality including stopovers;

d1 = B*x;
print(d1);

d2 = (B**2)*x;
print(d2);

d3 = (B**3)*x;
print(d3);

* as k--> infinity, the normalized vector v approaches the leading eigenvector of A and/or B;
v1 =  d1/sqrt(d1[##]);
print v1;

d5 = (B**5)*x;
print(d5);

v5 =  d5/sqrt(d5[##]);
print v5;

d10 = (B**10)*x;
print(d5);

v10 =  d10/sqrt(d10[##]);
print v10;

run; quit;

* Rick Wicklin's Power Method Function;
proc iml;
     start PowerMethod(v, A, maxIters);
   /* specify relative tolerance used for convergence */
   tolerance = 1e-6;
   v = v / sqrt( v[##] );  /* normalize */
   iteration = 0; lambdaOld = 0;

   do while ( iteration <= maxIters);
      z = A*v;                /* transform */
      v = z / sqrt( z[##] );  /* normalize */
      lambda = v` * z;
      iteration = iteration + 1;
      if abs((lambda - lambdaOld)/lambda) < tolerance then
         return ( lambda );
      lambdaOld = lambda;
   end;
   return ( . ); /* no convergence */
finish;

* try our data;
A = {0 1 1 1,
     1 0 1 0,
      1 1 0 0,
      1 0 0 0};

v = {1,1,1,1}; 

lambda = PowerMethod(v, A, 40 );
print lambda;
print v;

run;quit;