So, we've been talking about the ability for participants to publish a transaction and get it into the blockchain as if this happens by magic. Of course, it doesn't happen by magic in the real world, it happens through the Bitcoin network. So what is a Bitcoin network? It's a peer-to-peer network, so it inherits a lot of ideas from peer-to-peer networks that've been proposed for all sorts of other purposes, it's a peer-to-peer network where all nodes are equal. There's no hierarchy, there's no centralized special nodes, no master nodes, every node on Bitcoin is an equal peer. It runs over TCP, it has a random topology, so there's random nodes that appear with random other nodes, and new nodes can come at any time. So you can download the Bitcoin client today, you can spin your computer up as a node, and you'll be a participating Bitcoin node with equal rights and capabilities as every other node on the Bitcoin network. Now the network is very dynamic, it changes over time, nodes are coming and going all the time, although there's actually no explicit way to leave the network. Instead, if you don't hear from a node for a while, three hours is the amount that's hardcoded into the common clients, people eventually start to forget you, so it gracefully handles nodes going offline. So, what does that mean that you can simply join the Bitcoin peer-to-peer network at any time? Well, if this is a picture of the network at one moment in time, obviously scaled down quite a bit with just 7 nodes, but this is a picture of what it might look like. 7 nodes with all random connections to each other, and notice that the numbers are scattered around here because there's no geographic topology here, networks connect to other nodes in a random fashion by design. Now, if you launch a new node, and say you want to join the network you start with a simple message to one node that you know about, so all you need to know is how to get to one node that's already on the network. This is usually called your seed node, and there's a few different ways you can look up lists of seed nodes to try connecting to. But you find your seed node and you sent a special message saying tell me all the peers that you have, tell me the addresses of all the other nodes in the network that you know about. And that node will respond and say, well I'm peered with nodes 1 and 7, you can try them, and then you might go talk to 1 and 7 and say, hey tell me everybody on the network that you know about. And they'll send you the nodes that they know about and you can iterate as many times as you want until you have a list of peers to make connections with. And then you can choose which ones to peer with, and you'll be a fully functioning member of the Bitcoin network. And again there's a few steps of randomness here, so depending on which seed node you used or which of the peers of the seed node you decided to go and talk to, you'll end up with a random set of nodes that you're connected to, but that's perfectly fine. So now that you're a member of the network, what is the network good for? Well, the network maintains the blockchain, so if you want to publish a transaction you want to get the entire network to hear about it, and there's a simple flooding algorithm to make this happen. So let's say that Node 4 here, hears about a new transaction, so Alice wants to pay Bob some money, Alice creates a Bitcoin transaction and submits it to Node 4, or maybe her wallet software or her exchange does that on her behalf. But somehow this transaction gets to Node 4, now node 4 says great I've got a new transaction and Alice wants to pay Bob, let's tell everybody about it, sometimes this is called a gossip protocol because it's very simple. If you have news you try to tell as many people as you can and they try to tell as many people as they can, much like people gossiping in the real world. Great, so node 4 is going to talk to his neighbors node 3 and node 2 and say hey, check out this new transaction Alice wants to pay Bob, and those nodes will add it to their own pool of pending transactions. So, each node maintains a list of all the transactions they've heard about that haven't been put into the blockchain yet, and then they can decide to forward that onto other nodes. So 3 is going to talk to its neighbors and say, new transaction for you Alice wants to pay Bob, that'll end up in their transaction pools and so on. And we want to make sure that this process doesn't go on forever, so let's say that node 2 comes along later and tries to tell node 7, Hey, new transaction Alice wants to pay Bob, node 7 is going to say, That's all right node 2 I've already heard about that, I already got it in my memory pool I don't need to forward it further. So eventually, this thing has to stop, because every node will have heard about the new transaction and they won't forward it anymore. And remember, every transaction is identified uniquely by its hash, so each node can tell they've seen that hash before and that they don't need to keep forwarding that transaction, so it won't loop around the network forever. So how do nodes decide when they hear about a new transaction, whether or not they should propagate it? The most important thing they do Is they check to see, given their view of the block chain, whether or not this transaction is valid. So they do all the transaction validation we talked about earlier, they run the script, they see that the script checks out, they see that the coins are being redeemed here haven't already been spent, and if all of that checks out, then this looks like a valid transaction. That they should try to relay, with a couple of other caveats, by default nodes won't relay the transaction if it's a non-standard script. If the script has any weird features, if it doesn't match a fairly simple whitelist of scripts that nodes know about, even though it's a valid transaction, the nodes won't relay it. They'll also make sure that they haven't seen the transaction before, that's that condition to avoid infinite loops, and there's another property, which is that they won't relay the transaction if it looks like a double spend. So if they've seen a transaction where Alice tries to send some specific coins to Bob and then later they see a second transaction where Alice tries to send those same coins to Charlie, the node shouldn't relay the second transaction. Even though either transaction could be valid because those coins still haven't been spent, they'll only relay the first one they hear and that's an extra guard against double spending. But it's important to keep in mind that all of these checks are just sanity checks. So well behaving nodes, all implement these to try to keep the network healthy and running properly, but there's no rule that says that nodes have to follow these specific steps. So since it's a peer-to-peer network and anybody can join, there is always the possibility of a node not following this exact protocol, forwarding double spends, forwarding transactions that aren't standard, forwarding transactions that aren't valid. And that's why it's important that every node do the checking for itself. So it's possible that nodes will end up with a different view of the pending transaction pool based on what they've seen. So let's go back to this example where Node 4 originally relayed a transaction where Alice was trying to pay Bob. And let's say that this transaction hasn't yet flooded to the entire network. And before it gets to everybody node 1 is going to announce a new transaction and say hey, I just heard Alice is trying to pay Charlie. Now, from Node 1's perspective, this is a valid transaction, and they haven't seen the other transaction where Alice is trying to pay Bob. So, Node 1 is going to implement the protocol normally, and is going to tell all of her neighbors about it. Now the neighbors that haven't heard the conflicting transaction yet will add it to their transaction pool. Whereas other neighbors, like Node 6 in this example, they've already received a transaction where Alice is trying to pay Bob, so Node 6 is going to say I don't want to hold two conflicting transactions in my pool, I'll just keep the one I already have. The network may end up in divided state here where different nodes have a different view of what the pending transaction pool is but that's fine. These transactions haven't been published in the blockchain yet so this is just a temporary state where nodes disagree on which transactions should be put into the next block. In practice, this is a race condition. If nodes have a different perspective on which transactions are pending or which blocks have been accepted, that's okay in a temporary state, and eventually they'll sort it out. So in the case of transactions, if different nodes have a different view of the pending transaction pool, depending on who mines the next block, they'll essentially break the tie, or the race condition, and decide which of those two pending transactions should end up being put permanently into a block. And once one of those two transactions has been put into a block, other nodes will see that the transaction that they're holding onto in their pool Is now never going to make it into a block because it would be a double spend and they'll just drop it. So if the transaction where Alice tried to pay Bob, successfully makes it into a block first, the nodes who heard the transaction where Alice tried to pay Charlie will just say that's not a valid transaction anymore so I can forget it. So the default behavior is for nodes to just hang on to whatever they hear first, which means that network position matters. If two conflicting transactions or two conflicting blocks get announced at two different positions in the network, they'll both flood in opposite directions. And the nodes which end up with one transaction or the other will depend on which side of the network they started out closer to. Of course this assumes that every minor implements this logic where they keep whatever they hear first, but there's no central authority enforcing this, so every node is free to do whatever logic they want. So if for some reason, if anyone wants to, they can choose to implement any other logic they want for choosing which blocks or which transactions to forward. We'll talk about that more in our lecture on mining, why miners want to implement some different logic other than the default. Now I've been talking mostly about transactions here. The logic for announcing new blocks, whenever minors find a new block, is almost exactly the same as propagating a new transaction. So the same algorithm is used to announce new blocks around the network, it's the same flooding algorithm, the same gossip process. And in this case, instead of verifying that the transaction is valid by running a script, the nodes are going to verify that the new block is valid by computing the hash, and making sure that it starts with a sufficient number of zeroes to meet the difficulty target. Now validating a block is also much more in-depth. Because in addition to validating the header and seeing that the hash value is correct, nodes are asked to validate every transaction included in the block to make sure that the block contains only valid, new transactions. And the other check, which is this really important, critical part and that makes Bitcoin consensus what it is, is that nodes shouldn't forward a block unless it builds on their perspective of the current longest chain. So they have a view of the blockchain, and they should only forward new blocks if they come at the very end of the chain, not at some earlier point, and this avoids forks building up. So just like with transactions, nodes can implement different logic if they want, they're free to relay blocks that aren't valid or to relay blocks that build off of an earlier point in the block chain. So some nodes may be trying to relay a block that doesn't extend the current longest chain that actually builds a fork and that's okay, the protocol is designed to withstand that. So how long does this flooding algorithm actually take? How much latency is imposed here? This is a graph showing the average time for new blocks to propagate to every node in the network. And the three lines show the 25th, the 50th, and the 75th percentile. Of how long it takes for a new block to reach every node in the network. And if you look at the 75th percentile there for some of the larger blocks, and this is heavily dependent on size, because the bandwidth constraints that some nodes have, you'll see that the average propagation time is over 30 seconds. So this shows that this isn't a particularly efficient protocol. On the Internet, 30 seconds is a pretty long time for people to hear about something. The reason it takes so long is because the protocol is not very efficient, it wasn't designed to be efficient, it was designed to be simple and to have no structure, so that every node is equal and that they can come and go at every time. And as a result the topology may not be optimized for fast communication. A block may need to go through many nodes before it reaches some of the most distant nodes in the network. Whereas if you designed a network top down for efficiency, you would design it to make sure that the path between any two nodes was very short. For BitCoin, it's more important to have a decentralized structure where all nodes are equal, even if that means that the propagation time can be over 30 seconds in some cases. So how big is a BitCoin network? Well, there's no official statistics anywhere, because again, there's no central authority overseeing it, it's simply whatever the nodes participating they are the Bitcoin network. So it's impossible to measure exactly, and it's changing all the time, but a number of researchers that looked into this and tried to come up with estimates. On the high end, some researchers have said that over a million IP addresses in a given month will at some point be running the Bitcoin protocol and acting at least temporarily as a Bitcoin node. But if you look at full nodes that are actually permanently connected, and are fully validating every transaction they hear and running the full protocol, it's only about five or ten thousand, which may be a surprisingly low number. And in fact, that number may be dropping, there's no evidence that the number of fully validating nodes is going up, and there's some concern that the number of fully-validating nodes is actually going down. So to be a fully-validating node, you want to stay permanently connected so that you hear about all data. The longer you're offline the more catchup you're going to have to do to hear about all the transactions you missed, and you're going to have to store the entire blockchain. You'll also need a pretty active network connection so that you can hear every new transaction and forward it to your peers. So you can see the growth over time here, and currently it takes about 20 gigabytes to store the entire block chain. Which isn't too bad if you have a few years old PC with an active network connection, you have what it takes to be a fully validating node. Although you basically need to dedicate that machine to doing that and not much else. Fully validating nodes maintain the entire set of Unspent Transaction Outputs. So every coin that's available to be spent, and remember those are just unspent output transactions, ideally you'd like to store this in ram so that when you hear a new proposed transaction on the network you can quickly check the transaction that it's attempting to claim, run the script and see if the signature is valid. So currently there are about twelve million unspent transactions and that's out of 44 million transactions that have ever been proposed. So, fortunately that's still small enough to fit in less than a gigabyte of ram, in an efficient data structure. So that if you are running the fully validating node every time you hear about a new transaction, you can quickly check, run the redemption script and see that this is a valid transaction that you should put in your pending transaction pool. So in contrast to being a fully validating node, there are lightweight nodes, also called thin clients or simple payment verification clients. This is actually the vast majority of nodes on the BitCoin network, and the difference here is that these nodes aren't attempting to store the entire blockchain. They only store the pieces that they need to verify some specific transactions that they care about. So for example if you run a wallet, your wallet might want to be a simple payment verification node, and if somebody sends money to you, you'll act as a node. You'll download the bits of the blockchain that you need to verify that the person sending you the money actually owned it and that the transaction sending it to you actually gets included in the block chain. You won't care about the thousands of other transactions going on that don't affect you. Now an SPV client like this won't have the full security level of being a fully validating node. And the reason is that when they hear a new block, the only thing they can check is the block header. They can check to see that the block was difficult to mine, but they can't check to see that every transaction included in that block is actually valid. Because they don't have the entire previous blockchain, they don't know the entire unspent transaction output set. They can only validate the transactions that actually affect them. So, they're essentially trusting the fully validating nodes to have validated all the other transactions that are out there. So, this isn't a bad security trade off, you're assuming there are fully validating nodes out there that are doing the hard work. And then if miners went through the trouble to mine this block, which is a really expensive process, they probably also did some validation to make sure that this block wouldn't be rejected. And the cost savings of being an SPV node are huge. It's about a thousand times smaller to just store block headers than to store all of the previous transactions. Instead of storing about 20 gigabytes of data, you're down to about 20 megabytes. Which is something that almost anybody on a PC or even on a phone can store and act as a limited node in the BitCoin network.