diff options
Diffstat (limited to 'sourcecodes/BNW_workflow_net1.htm')
| -rw-r--r-- | sourcecodes/BNW_workflow_net1.htm | 49 |
1 files changed, 20 insertions, 29 deletions
diff --git a/sourcecodes/BNW_workflow_net1.htm b/sourcecodes/BNW_workflow_net1.htm index aafab31f..3ad19201 100644 --- a/sourcecodes/BNW_workflow_net1.htm +++ b/sourcecodes/BNW_workflow_net1.htm @@ -294,7 +294,7 @@ ul line-height:115%;font-family:"Arial","sans-serif"'> <br> This tutorial provides -an overview of using BNW to build a Bayesian network model from a dataset and use the network to make predictions. The dataset used in this tutorial is a synthetic example of a genetic dataset that has a total of 8 variables. Two of the variables are genotypes labeled Geno1 and Geno2, and the remaining 6 variables are gene expression levels or other quantitative traits that are labeled Trait1 to Trait6. The dataset is available <a href="example_datasets/example_data_8nodes.txt">here</a>.<br><br>The data file is formatted according to the guidelines on the <a href=http://compbio.uthsc.edu/BNW_1.2/sourcecodes/help.php#file_format>BNW help page</a>. The first row of the file contains the names of the variables and the remaining rows contain the data for each sample of the dataset. The genotypes (Geno1 and Geno2), which have two possible states (1 and 2), are the only discrete variables in the network and are the leftmost variables in the input file. The quantitative traits are continuous variables.<o:p></o:p></span></p> +an overview of using BNW to build a Bayesian network model from a dataset and use the network to make predictions. The dataset used in this tutorial is a synthetic example of a genetic dataset that has a total of 8 variables. Two of the variables are genotypes labeled Geno1 and Geno2, and the remaining 6 variables are gene expression levels or other quantitative traits that are labeled Trait1 to Trait6. The dataset is available <a href="example_datasets/example_data_8nodes.txt">here</a>.<br><br>The data file is formatted according to the guidelines on the <a href=./help.php#file_format>BNW help page</a>. The first row of the file contains the names of the variables and the remaining rows contain the data for each sample of the dataset. The genotypes (Geno1 and Geno2), which have two possible states (1 and 2), are the only discrete variables in the network. The quantitative traits are continuous variables.<o:p></o:p></span></p> <p class=MsoNormal style='margin-right:107.5pt'><b style='mso-bidi-font-weight: normal'><span style='font-size:14.0pt;line-height:115%;font-family:"Arial","sans-serif"'>1. Structure learning using default options<o:p></o:p></span></b></p> @@ -303,17 +303,17 @@ normal'><span style='font-size:14.0pt;line-height:115%;font-family:"Arial","sans 0in;margin-left:0in;margin-bottom:.0001pt'><span style='font-size:12.0pt; line-height:115%;font-family:"Arial","sans-serif"'>We do not know the network structure that underlies the relationships between the variables in this dataset, so we will use BNW to learn the structure that best explains the data. To begin, select <u>Learn a network model from data</u> from the BNW home page. Next, click <u>Choose File</u> at the top of the file upload page, navigate to and select the data file, and click <u>Upload</u>. A screen similar to image shown below should be displayed.<br><br><o:p></o:p></span></p> <p class=MsoNormal style='margin-right:107.5pt'><span style='font-size:12.0pt; -line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'><![if !vml]><img width=1128 height=499 -src="BNW_workflow_test_files/8node_file_upload.jpg" v:shapes="Picture_x0020_2"><![endif]></span><span +line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'><![if !vml]><img width=787 height=423 +src="BNW_workflow_test_files/8node_file_upload_new.png" v:shapes="Picture_x0020_2"><![endif]></span><span style='font-size:12.0pt;line-height:115%;font-family:"Arial","sans-serif"'><o:p></o:p></span></p> <p class=MsoNormal style='margin-top:0in;margin-right:107.5pt;margin-bottom: 0in;margin-left:0in;margin-bottom:.0001pt'><span style='font-size:12.0pt; -line-height:115%;font-family:"Arial","sans-serif"'>Clicking on the <u>View uploaded variables and data</u> on the left menu will bring up a pop-up window that displays the uploaded data file or information about the data set, such whether variables were classified as discrete or continuous. This information can be used to ensure that the input file was uploaded and properly interpreted in BNW.<br><br> Other options on the left menu will continue with structure learning. Initially, we will perform structure learning using default options in BNW and click on the <u>Perform Bayesian network modeling using default settings</u> button. By default, BNW limits the maximum number of parents for any node in the network to 4 and presents the structure of the single highest scoring network. Structure learning can take a significant amount of time for larger networks. For this dataset, structure learning should take only a few seconds, and the structure below will soon be displayed. The network can be also be accessed <a href="example.php?My_key=example1|Bqx" target="_blank">here</a>.<br><br> +line-height:115%;font-family:"Arial","sans-serif"'>The buttons on the top of the page allow you to perform Bayesian network modeling using default settings; modify structure learning settings and add structural constraints; or remove variables from the data set. Under these buttons, a description of the loaded data set is provided, including if BNW classified the variables as discrete or continuous. In this case, the uploaded data file contained 8 variables and had 200 cases. Two variables (Geno1 and Geno2) were determined to be discrete variables because they had only 2 different values in the input data file. The other six variables were continuous variables with the means and standard deviations shown in the table. <br><br> Initially, we will perform structure learning using default options in BNW and click on the <u>Perform Bayesian network modeling using default settings</u> button. By default, BNW limits the maximum number of parents for any node in the network to 4 and presents the structure of the single highest scoring network. Structure learning can take a significant amount of time for larger networks. For this dataset, structure learning should take only a few seconds, and the structure below will soon be displayed. The network can be also be accessed <a href="example.php?My_key=example1|Bqx" target="_blank">here</a>.<br><br> <p class=MsoNormal style='margin-right:107.5pt'><span style='font-size:12.0pt; line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'> -<img width=908 height=842 -src="BNW_workflow_test_files/8node_best_structure.jpg" v:shapes="Picture_x0020_5"><o:p></o:p></span></p> +<img width=308 height=264 +src="BNW_workflow_test_files/8node_no_weight.png" v:shapes="Picture_x0020_5"><o:p></o:p></span></p> <p class=MsoNormal style='margin-top:0in;margin-right:107.5pt;margin-bottom: 0in;margin-left:0in;margin-bottom:.0001pt'><span style='font-size:12.0pt; @@ -324,58 +324,49 @@ normal'><span style='font-size:14.0pt;line-height:115%;font-family:"Arial","sans <p class=MsoNormal style='margin-top:0in;margin-right:107.5pt;margin-bottom: 0in;margin-left:0in;margin-bottom:.0001pt'><span style='font-size:12.0pt; -line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'>In order to test if edges present the single best scoring network are conserved across high scoring networks. We can modify the structure learning settings to get identify the structures of many high scoring networks and perform <a href=http://compbio.uthsc.edu/BNW_1.2/sourcecodes/help.php#learn_details>model averaging</a> over these structures. To do this, return to the BNW home page, select <u>Learn a network model from data</u>, and upload the data file. Instead of using the default settings, select <u>Go to structure learning settings and the BNW structural constraint interface</u>. A more detailed overview of use of the structural constraint interface is provided in <a href=http://compbio.uthsc.edu/BNW_1.2/sourcecodes/BNW_workflow_2.htm>another tutorial</a>, but, here, we will investigate the impact of modifying some of the structure learning settings shown below:<br><o:p></o:p></span></p><br> +line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'>In order to test if edges present the single best scoring network are conserved across high scoring networks, we can modify the structure learning settings to identify the structures of many high scoring networks and perform <a href=./help.php#learn_details>model averaging</a> over these structures. To do this, click the <u>Modify network strucutre</u> button on the left of the page, opening a menu with three options that can be used to modify the network structure. <u>Add or remove edges from network</u> provides an interface for making specific changes to the network structure. The use of this interface is discussed <a href=./add_remove_howto.htm target=_blank>here</a>. In this tutorial, we are examining the impact of changing the settings used during structure learning on the network structure and click on <u>Modify structure learning settings</u>. We then select <u>Go to structure learning settings and the BNW structural constraint interface</u> on the resulting page. A more detailed overview of use of the structural constraint interface is provided in <a href=./BNW_workflow_sci.htm>another tutorial</a>. <br><br>Here, we only change the <u>Number of networks to include in model averaging</u> to 1000 and select <u>Continue to assign variables to tiers</u>. In this case, we will not assign variables to tiers, and can immediately click <u>Click here to perform Bayesian network modeling after creating tiers</u>. Now, instead of displaying the single highest scoring network, BNW will determine the 1000 highest scoring networks, perform model averaing over these networks, and display the structure after model averaging that includes all features with a <u>Model averaging edge selection threshold</u> greater than 0.5. Model averaging over the 1000 highest scoring structures has resulted in a change in the network structure as shown below and is available <a href="example.php?My_key=example2|hQG" target="blank_">here</a>.<br><o:p></o:p></span></p><br> <p class=MsoNormal style='margin-right:107.5pt'><span style='font-size:12.0pt; line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'> -<![if !vml]><img width=628 height=181 -src="BNW_workflow_test_files/8node_global_settings.jpg" v:shapes="Picture_x0020_8"><![endif]><o:p></o:p></span></p> +<![if !vml]><img width=366 height=262 +src="BNW_workflow_test_files/8node_model_avg_struct_new.png" v:shapes="Picture_x0020_8"><![endif]><o:p></o:p></span></p> <p class=MsoNormal style='margin-top:0in;margin-right:107.5pt;margin-bottom: 0in;margin-left:0in;margin-bottom:.0001pt'><span style='font-size:12.0pt; -line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'><br>Change the <u>Number of networks to include in model averaging</u> to 100 and select <u>Perform Bayesian network modeling</u>. Now, instead of displaying the single highest scoring network, BNW will determine the 100 highest scoring networks, perform model averaing over these networks, and display the structure after model averaging that includes all features with a <u>Model averaging edge selection threshold</u> greater than 0.5. Model averaging over the 100 highest scoring structures has resulted in a change in the network structure as shown below and is available <a href="example.php?My_key=example2|hQG" target="blank_">here</a>.<br><o:p></o:p></span></p><br> +line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'><br>Specifically, the single best model network contains a directed edge linking Trait6 with Trait3, while this edge is absent from the structure after model averaging over the 1000 best scoring networks. More information about the network structure can be viewed by clicking <u>Network image options</u> on the left menu and, then, <u>Show network with edge weights</u>. The network structure with the edges labeled by their scores after model averaging will be displayed.<br><o:p></o:p></span></p><br> <p class=MsoNormal style='margin-right:107.5pt'><span style='font-size:12.0pt; line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'> -<![if !vml]><img width=1034 height=802 -src="BNW_workflow_test_files/8node_model_avg_struct.jpg" v:shapes="Picture_x0020_8"><![endif]><o:p></o:p></span></p> +<![if !vml]><img width=389 height=259 +src="BNW_workflow_test_files/8node_with_weights.png" v:shapes="Picture_x0020_8"><![endif]><o:p></o:p></span></p> <p class=MsoNormal style='margin-top:0in;margin-right:107.5pt;margin-bottom: 0in;margin-left:0in;margin-bottom:.0001pt'><span style='font-size:12.0pt; -line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'><br>Specifically, the single best model network contains a directed edge linking Trait6 with Trait3, while this edge is absent from the structure after model averaging over the 100 best scoring networks. Clicking <u>Display structure matrix</u> displays the model averaging scores as well as the structure matrix. The structure matrix file can be downloaded and used in return sessions to BNW, allowing users to skip structure learning and more quickly use the model to make predictions.<br><o:p></o:p></span></p><br> - -<p class=MsoNormal style='margin-right:107.5pt'><span style='font-size:12.0pt; -line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'> -<![if !vml]><img width=899 height=563 -src="BNW_workflow_test_files/8node_matrix.jpg" v:shapes="Picture_x0020_8"><![endif]><o:p></o:p></span></p> - -<p class=MsoNormal style='margin-top:0in;margin-right:107.5pt;margin-bottom: -0in;margin-left:0in;margin-bottom:.0001pt'><span style='font-size:12.0pt; -line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'><br>In the model averaging scores, we can observe that most of the network edges in the model average network were found in all or nearly all of the 100 highest scoring networks. For example, edges from Geno2 to Traits 3 and 4 have posterior probabilities of 1. Also, the edge from Trait6 to Trait3 that was observed in the single highest scoring network has a posterior probability of 0.46, and it, therefore, falls just short of being included in the model averaged network structure.<br><o:p></o:p></span></p><br> +line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'><br> The model averaging scores show that the edges in this network model were found in all or nearly all of the 1000 highest scoring networks, and, thus, have scores of 0.99 or 1.00. For example, edges from Geno2 to Traits 3 and 4 have scores of 1. The model averaging scores for all possible directed edges in the network can be found by clicking <u>View structure matrix</u> under the <u>More about network</u> menu. The data provided in the structure matrix shows that the edge from Trait6 to Trait3 that was observed in the single highest scoring network has a posterior probability of 0.40, and it, therefore, falls short of being included in the model averaged network structure. The structure matrix file can be downloaded and used in return sessions to BNW, allowing users to skip structure learning and more quickly use the model to make predictions.<o:p></o:p></span></p><br> <p class=MsoNormal style='margin-right:107.5pt'><b style='mso-bidi-font-weight: normal'><span style='font-size:14.0pt;line-height:115%;font-family:"Arial","sans-serif"'>3. Using the network structure to make predictions<o:p></o:p></span></b></p> <p class=MsoNormal style='margin-top:0in;margin-right:107.5pt;margin-bottom: 0in;margin-left:0in;margin-bottom:.0001pt'><span style='font-size:12.0pt; -line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'>To make predictions with the network, we will use the structure learned after model averaging of the top 100 highest scoring networks. First, we will use the model to compare the expected values for nodes in the network based on observed genotypes. For these predictions, we will use evidence mode when making predictions, which is the default behavior in BNW. The difference between evidence and intervention modes is discussed in the <a href=http://compbio.uthsc.edu/BNW_1.2/sourcecodes/faq.php#evid_inter>BNW FAQ page</a>. To use the model to make predictions based on Geno1, click on one of the blue bars in the Geno1 node and enter 1 or 2 to indicate which genotype value should be used to predict the values of the other network nodes. In the figure below, Geno1 is outlined in red and state 2 has a 100% probability, indicating that the value of this node has been entered as evidence. The red lines in the figure show the predicted distributions of the nodes after this evidence is known and can be compared with the blue lines which show the distributions for variables using the original data.<o:p></o:p></span></p><br> +line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'>The network model can be used by make predictions through an interactive interface after clicking on the <u>Use network to make predictions</u> button on the left menu. To make predictions with the network, we will use the structure learned after model averaging of the top 1000 highest scoring networks. First, we will use the model to compare the expected values for nodes in the network based on observed genotypes. For these predictions, we will use evidence mode when making predictions, which is the default behavior in BNW. The difference between evidence and intervention modes is discussed in the <a href=./faq.php#evid_inter>BNW FAQ page</a>. To use the model to make predictions based on Geno1, click on one of the blue bars in the Geno1 node and enter 1 or 2 to indicate which genotype value should be used to predict the values of the other network nodes. In the figure below, Geno1 is outlined in red and state 2 has a 100% probability, indicating that the value of this node has been entered as evidence. The red lines in the figure show the predicted distributions of the nodes after this evidence is known and can be compared with the blue lines which show the distributions for variables using the original data.<o:p></o:p></span></p><br> <p class=MsoNormal style='margin-right:107.5pt'><span style='font-size:12.0pt; line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'> -<![if !vml]><img width=1086 height=866 -src="BNW_workflow_test_files/8node_geno1_2.jpg" v:shapes="Picture_x0020_8"><![endif]><o:p></o:p></span></p> +<![if !vml]><img width=533 height=517 +src="BNW_workflow_test_files/8node_geno1_2_new.png" v:shapes="Picture_x0020_8"><![endif]><o:p></o:p></span></p> <p class=MsoNormal style='margin-top:0in;margin-right:107.5pt;margin-bottom: 0in;margin-left:0in;margin-bottom:.0001pt'><span style='font-size:12.0pt; -line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'><br>If Geno1 has genotype 2, the value of Traits 1, 2, and 4 are expected to increase compared with the distribution for all data. Specifically, the mean value of Trait2 is expected to be near 1 for Geno1=2 data, while it is close to 0 when this evidence is not known. Traits 1 and 4 are also expected to increase, but the magnitude of this increase is not expected to be as large. Predicted distributions for the other nodes in the network, which are not descendants of Geno1, are expected to be close to the same as their original distributions, and, the red line covers the blue line for some nodes.<br><br> -To quantitatively assess the impact of this evidence on the network predictions, the <u>View parameters</u> button on the left menu can be selected. Clicking this button brings up a pop-up window with the network parameters (i.e., the probability distributions of the states of discrete nodes and the means and standard deviations of the Gaussian distributions for continuous nodes) for both the original data set and the data when considering the entered evidence. <br><br> +line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'><br>If Geno1 has genotype 2, the value of Traits 1, 2, and 4 are expected to increase compared with the distribution for all data. Specifically, the mean value of Trait2 is expected to be near 1 for Geno1=2 data, while it is close to 0 when this evidence is not known. Traits 1 and 4 are also expected to increase, but the magnitude of this increase is not expected to be as large.<br><br> +To quantitatively assess the impact of this evidence on the network predictions, the <u>View parameters</u> button within the <u>More about network parameters and structure</u> menu can be selected. Clicking this button brings up a pop-up window with the network parameters (i.e., the probability distributions of the states of discrete nodes and the means and standard deviations of the Gaussian distributions for continuous nodes) for both the original data set and the data when considering the entered evidence. <br><br> Evidence for multiple nodes can be considered at the same time by clicking on a new node in the network and entering a value. Alternatively, users can select <u>Clear evidence</u> to reset the network to show the orignial distributions in a new tab.<br><br> Next, we will make predictions using the intervention mode. To use the prediction mode, click the button next to <u>Intervention</u> at the top of the page. The <u>Selected mode</u> tab on the left of the screan should now display Intervention. The effects of experimental intervention on Trait2 can be predicted by clicking on the blue line in the Trait2 node and entering a value for the variable. The figure below shows the network after setting Trait2 to a value of 1.5.<o:p></o:p></span></p><br> <p class=MsoNormal style='margin-right:107.5pt'><span style='font-size:12.0pt; line-height:115%;font-family:"Arial","sans-serif";mso-no-proof:yes'> -<![if !vml]><img width=1036 height=864 -src="BNW_workflow_test_files/8node_trait2_15.jpg" v:shapes="Picture_x0020_8"><![endif]><o:p></o:p></span></p> +<![if !vml]><img width=517 height=516 +src="BNW_workflow_test_files/8node_trait2_15_new.png" v:shapes="Picture_x0020_8"><![endif]><o:p></o:p></span></p> <p class=MsoNormal style='margin-top:0in;margin-right:107.5pt;margin-bottom: 0in;margin-left:0in;margin-bottom:.0001pt'><span style='font-size:12.0pt; |
