Showing posts with label UCS. Show all posts
Showing posts with label UCS. Show all posts

Monday, August 22, 2011

Let's Talk about Cisco UC on Vblock at VMworld!

For a good portion of VMworld this year I have the distinct privilege of hosting customers in the VCE Booth (#1121).  We have many great things to discuss this year that I will be highlighting over the next few days leading up to VMworld.

Today, I want to discuss a solution near and dear to my heart.  My team has been working closely with Cisco to bring Cisco Unified Communications (UC) as an integrated solution to the Vblock platform.

What has VCE accomplished to date for Cisco UC?
Through VCE and Cisco's collaboration, we have introduced Unified Communications Blade Packs specific to the Vblock platform.  If you are familiar with UC, you will understand the concept of Cisco's Tested Reference Configurations.  Basically, it is a hardware standard that the Cisco UC team has fully tested and Cisco will support.  Cisco and VCE jointly support the B200 M2 TRC#1 and TRC#2 configurations.  In addition, EMC's PowerPath/VE and the Cisco 1000v now come standard with every UC blade pack to better support co-residency with other applications inside a Vblock (but on different hardware blades).  If you have any questions, find me at VMworld and we'll discuss the "nitty gritty" details.

Why has VCE introduced the Cisco UC Blade Packs?
They say a picture is worth a thousand words.  The following table is taken from the design considerations page on Cisco's UC website.  We want to make the infrastructure "go away".  By running UC on Vblock, you can focus on the UC design and not worry about all the "stuff underneath".  All components of the Vblock are supported by VCE & Cisco will handle the UC application.  It couldn't be easier!





What about UC's newly announced specs-based support?
In June of this year Cisco introduced the concept of specs-based support for UC.  Long term this will remove the hardware dependence of the Tested Reference Configurations.  VCE and Cisco are working closely on a fully supported and documented solution for this architecture as well.  Specs-based UC on Vblock will work today but it is up to the customer or partner to design the architecture.

In conclusion, it's an exciting time for Cisco UC on Vblock!  Stop by VCE's Booth 1121 and let's talk about it!

Wednesday, July 20, 2011

Scale Up with VMware vSphere 5: “I’m Not Dead Yet!”

Well, what a crazy two weeks. It looks like Kevin Bacon became a VMware employee this week:

Since this post will be long and filled with numbers, here’s a summary:

The cost difference to license VMware vSphere 5 in a Scale Out vs. Scale Up scenario is roughly equivalent, BUT if you take into account the additional hardware and Microsoft license costs required, Scale Up holds it own against Scale Out. Remember, we have to build a TOTAL solution, not just look at one small piece of the overall cost.

Now that the panic has died down a bit, I wanted to show some very surprising numbers that I found last week regarding Scale Up vs. Scale Out with vSphere 5. Many people’s first reaction to high memory environments was “I’m just going to buy 96GB servers from now on!” I’ve run the numbers and the short answer is you shouldn’t do that if your main concern is a higher TCA (Total Cost of Acquisition).

For those unfamiliar with the terms, what are Scale Out and Scale Up? Scale Out refers to the concept of hosting many small servers in a cluster to provide resources. For this article, I will define a 2xCPU server with 96GB of memory as the baseline for Scale Out. In Scale Out, you have many a higher number of smaller sized servers. The main advantage to Scale Out is a higher utilization per node if you plan for an outage in the environment. Today we use n-1 or n-2 typically for planning purposes. This means we set aside the capacity equal to one or two hosts spread across the cluster. Scale Out also allows you to start small and grow large over time.

Scale Up is the concept of using a smaller number of servers with a larger capacity to achieve that same cluster size of resources. In this instance, you have a smaller number of larger sized servers.  For this article I classify a Scale Up server as anything greater than 2xCPU or 96GB of memory. The advantage to Scale Up is fewer management points and typically a smaller cost per unit. In a Scale Up scenario, you invest up front and see greater savings over time.

I know many customers have different theories with either model and no one model fits every customer. If you take into account budget cycles, internal politics, facilities, existing standards, etc. you may not be able to make a decision based solely on the numbers presented here. I also concede that some just aren’t comfortable with Scale Up and the potential virtual machine density it can provide. In Scale Up’s defense, I contend that the advances to vMotion and HA in both vSphere 4.1 and 5 we will see this fear start to ease over time as everyone becomes more comfortable with the enhancements to the technologies.

If VMware license costs aren’t the Scale Out culprit, who is?

Everyone get your pitchforks; let’s figure out who we need to string up next! As I stated previously, it is now cost neutral for Scale Out and Scale Up with vSphere. But, what about Microsoft licenses? What about hardware costs for new servers? What about soft facility costs like power, cooling and space? If you are going to build a solution you need to factor in ALL the costs, not just a single data point.

Let’s take each item one by one.
  • Microsoft License Costs – Uncle Bill Gates (or should I say Uncle Steve Balmer these days) gets paid one way or the other. In larger environments, the most cost effective way today to license a VMware server is purchase a Data Center License for each host. A Microsoft Data Center License costs $2999 MSRP per socket. This leads to a minimum $6000 “mTax” per server! 
  • Physical Server Costs – There are many things to consider here. We need to buy a server, at least 96GB of memory, 10GB NICs or a bunch of 1GB NICs, and we can’t forget about hardware maintenance for 3-5 years. I went online to the HP website and configured a 2xCPU, 96GB, 2x10GB NICs DL360 for about $11,000 MSRP. I then calculated the same server with 192GB for about $24,000 MSRP. I’m going to use these two numbers as my data points. Let’s call this value the “pTax”. 
  • Soft Costs – This includes many things that are hard to quantify but NEED to be considered in calculations. It costs money to own and operate a server. They take power, they take cooling, they need people to manage them and we have to pay this staff (at least the good ones), they consume network ports, we have to plug cables into them, etc. Every thing listed contributes to the cost of the solution. Let’s call this value the “sTax”. This isn’t a fantasy world of free servers and facilities. No Rainbows and Unicorns for You!

Still with me? Eyes haven’t glazed over yet? In for the long haul? Ready for some numbers? Read On.

I created a spreadsheet modeled after the awesome Rynardt Spies’s Blog Post: http://www.virtualvcp.com/news/163-a-deeper-look-into-vmware-vsphere-5-licensing

For this exercise, I’m assuming only 2xCPU servers per the pTax bullet above and comparing 96GB and 192GB configurations. I’m also assuming vSphere Enterprise Plus and MSFT Data Canter Licenses for each server. This is a worst-case scenario: a scenario where the entire infrastructure is purchased up front. In environments of this size, it is often considered bad practice to “go back to the well” for more funding later. I have consulted with many customers in which this has been the case over the years.

Let’s look at a small four host, Scale Up 192GB per server, cluster vs. an equivalent amount of 96GB servers. To calculate the number of equivalent 96GB Scale Out servers, we need to compare the amount of the Configured Memory value (highlighted below). Using an n-1 failure scenario, you can see if you take the amount of cluster memory (768 GB) on the 192GB line and then subtract out the memory needed to support an n-1 host failure (3 hosts * 192GB = 576) you can achieve an equivalent amount of configured memory utilizing Scale Out 96GB with seven hosts.

Now that we know the number of hosts for Scale Out and Scale Up, we can calculate the total cost. In this example 3 subcomponents, vSphere cost, MS DC cost, and hardware cost, provide the total cost. vSphere cost is equal to the number of licenses * $3495, MS DC costs is equal to the number of licenses * $2999, and hardware is equal to the number of hosts * either $24,000 for Scale Up or $11,000 for Scale Out.

As you can see in this example, Scale Out is marginally more expensive (I’d argue that the cost is close enough to be “in the noise” and the costs are roughly equal). Scale Up is a viable solution.


What happens when we double the cluster size to 8 Scale Up hosts? Using the same calculations we see Scale Up start to pull ahead.


Let’s go big! In this example I doubled again to 16 Scale Up hosts. In this example I used n-2 for the host failure scenario due to the increased cluster size. The greater we scale, the greater the benefits.


In conclusion, Scale Up as a general rule isn’t more expensive. As memory costs fall, this trend will accelerate over time. Scale Up is often equal too or less than the TCA (Total Cost of Acquisition) for an equivalent amount of Scale Out computing. Scale Up can provide a beneficial cost structure while providing less management points in your environment. Remember, if you are going to include the “vTax”, you’ll need to include the “mTax” “pTax” and “sTax” as well!

Scale Out Computing’s Not Dead Yet!


I want to thank Maish, Andrew Storrs & vTexan as well as some of my VCE & VMware peeps for helping me out with the numbers and providing clarity. This post was a train wreck to start would not have been possible without them!!!

Wednesday, April 27, 2011

New Best Practice for Creating an ESXi Default Scratch Partition?

Some interesting blog posts came about yesterday with the start of this article on ESXi Scratch Partition Best Practices on the VMware ESXi Chronicles Blog.  This VMware KB article includes more technical detail as well as a resolution.  In the KB article, it states that many SAS and Boot From SAN(BFS) installations will NOT create a default scratch partition at install due to the possible shared nature of both the SAS and BFS architectures.

Both Scott Lowe and Forbes Guthrie followed up with articles based on previous ESXi installation experience.  Their great articles started me thinking one step further.  Since I tend to think of servers in terms of Cisco UCS these days, doesn't this mean that ALL Cisco UCS (and most other vendor servers as well) will not have a default scratch space because most installations are now either SAS or BFS??  If this is the case, shouldn't this KB article represent a new best practice for ALL installations of ESXi in the future as well as all installations in the past??


UPDATE:  Jeremy Waldrop shared with me that many of his UCS BFS installs include the scratch space by default and Scott Lowe shared on his blog that sometimes his didn't.  Looks like we have a "feature" on our hands where sometimes the scratch partition is created and sometimes it isn't.  My recommendation to everyone is to make sure you check for the partition post installation until we gather more information on the subject.


Does this need to be a best practice?  I'm sure the answer here will be "it depends."  It depends on how much of a difference not having the scratch space is to you.  Read the KB article carefully and understand what you are losing by not having a default scratch partition.  From reading the article, it seems worth it to me.  What are your thoughts?

Monday, March 21, 2011

Updates to Cisco UC on UCS Blades

Cisco recently announced a pretty big change to the support of Cisco Unified Communications (UC) products on the UCS B-series blade servers.  For select UC applications, the following features have been added to the support statement:

  • VMware ESXi 4.1 is now a supported hypervisor
  • Boot from FC is now supported for ESXi and a diskless configuration has been tested
  • limited vMotion is now supported but VMware DRS is not (i.e. you can manually move machines around if you have a need but VMware won't load balance them for you)
  • Cloning of vm's (clone at the vm level and then CUCM CLI) now allows template based installations to accelerate provisioning
All of the details can be found at the Cisco UC on Virtualization Wiki.  As the Cisco UC support statements are updated, I will post updates here.

Oh, and by the way, Cisco UC is fully supported on VCE Vblocks today...

Thursday, February 3, 2011

Links to Great Cisco UCS and Nexus 1000v Links

I'm a bit behind on some of these links but I wanted to present some great Cisco UCS and Nexus 1kv links from the past few weeks.


Wednesday, June 9, 2010

Comparing Vblocks

I believe one of the most interesting concepts to come along in our industry recently has been Cisco/EMC/VMware's Vblock.  My best definition for Vblock is a reference architecture that you can purchase.  Think about that for a second.  Many vendors publish reference architectures that are guidelines for you to build to their specifications.  Vblock is different because it is a reference architecture you can purchase.  This concept is a fundamental shift in our market to simplify the complexity of solutions as we consolidate Data Center technologies.  We are no longer purchasing pieces and parts, we are purchasing solutions.
Anybody who knows me knows I love the term "cookie cutter".  I use it all the time because it very simply conveys the idea of mass replication in a predefined way.  Vblock is a "Cookie Cutter" Data Center.  As long as you stay within the guidelines presented in the reference architecture, the product is guaranteed to work.
I took some time this week to compare and contrast all of the various Vblock configurations.  Take particular notice to the items highlighted in yellow below.  Vblock 0 is very different from Vblock 1 and 2 due to the basis in IP storage vs. FC storage as well as using ESXi vs. ESX.  Here are my findings (Please excuse the use of a graphic for the table, blogger sucks at tables):


UPDATE (I forgot this part) - A Few Notes on Vblock0
Because Vblock0 boots from local disks instead of FC boot from SAN, you lose the stateless ability of Cisco UCS.  This isn't a big deal in a smaller environment but stateless is increasingly important as the solution scales up. The Vblock0 reference document doesn't list the disk characteristics at all.  As a matter of fact, it makes no mention of the disk configuration and the IOP's Block 0 will generate.  I hope the document is updated to include this information in the near future.


Now, the million dollar question: How many virtual machines can I fit on each solution?  The reference materials I used for this didn't really provide numbers that I would use so I have decided to use my own.  My assumptions are presented as well as my work in case you want to change the math, or call me an idiot.
The Vblock 1 and Vlock 2 Reference Architecture document lists a 4 vm's per core estimate using 2GB for the virtual machine size so I thought I would start there.  vBlock 0 and vBlock 1 contain blades with 48GB. vBlock 1 also contains a few blades with 96GB and vBlock 2 is all 96GB.  Here is the math I used to figure out the proper ratio of vm's per core per GB:

4 vm's per core w/ 48GB:
4 vm's per core * 8 cores per blade = 32 vm's per blade * 2 GB per vm = 64 GB total needed / 48 GB total memory = 1.3 oversubscribe (30% memory over subscription)

Using the numbers presented in the reference architecture works out pretty well.  I wouldn't be comfortable with a higher density than that for the 48GB solution.

8 vm's per core w/ 96GB:
8 vm's per core * 8 cores per blade = 64 vm's per blade * 2 GB per vm = 128 GB total needed / 96 GB total memory = 1.3 oversubscribe (30% memory over subscription)

Since the blade has double the memory, I doubled the vm's per core to achieve the same over subscription ratio.  I don't mind telling you that 64 vm's on a blade takes me to the edge of my comfort zone but I'll accept that density level for this calculation.

Using the values of 32 vm's per 48GB blade and 64 vm's per 96GB blade you achieve the following minimum and maximum ranges for each of the Vblocks.  Remember the Vblock 1 contains BOTH 48GB and 96GB blades so the math gets a little harder there.

Aaron's Fuzzy Math Vblock Minimums and Maximums:

There you have it.  What do you think?  Am I even close?  Please leave a comment!

Wednesday, May 5, 2010

Verifying FEX Uplink Pinning in Cisco UCS

In the past I covered how to design for the necessary number of uplinks between the FEX's (IOM's) and the 6100's on Cisco UCS.  Once you have calculated the number of uplinks you will need, how do you know in UCS that the blades have pinned to the uplinks correctly?  (If you don't know what this means, please see my FEX design article listed above, pinning is covered there).  The short answer is that there isn't anything in the UCSM GUI today to tell you this.  The closest thing to a pin mapping output is a listing of the number of uplinks and no errors in the UCSM as an indication of correct pinning.  Maybe I'm too "detail oriented" but I like more information that.  So, off to the command line.

You will need the following commands to verify the status of pinning in UCS.

Step 1 - Connect to the nxos context in the 6100 in question by submitting the command: connect nxos


Step 2 - The command show fex will display the FEX's currently connected to the 6100.  In this case I have two UCS chassis:

Step 3 - Use the command show fex (fex#) details to output the actual pinning relationship.  Empty blade slots are not shown at all and powered off blades are shown in a down status.  In the first example I have 4 FEX uplinks, blades present in slots 1-4, and power applied to blades 1 & 4.  In the second example I have 1 FEX uplink, blades in slots 5-8, and all blades are turned on.

There you have it.  Now you can tell for sure that your blade pinning is correct and this process is also very helpful for troubleshooting as well.

Sunday, May 2, 2010

Cisco UCS Pools, Policies, Profiles and Templates

As I discussed in my previous post on UCS Pools, a service profile is the glue that holds Cisco UCS together at the server level.  Without service profiles being associated to the hardware, the blades won't boot!  As with everything else UCS, there is an high level of customization power around service profiles.  The other day I mentioned on Twitter that there is an amazing number of knobs you can turn.  Think of Cisco UCS something like this:

How do make sense of all of this?  As tends to be my way, I'll start with some definitions and then glue them together into the concept.

Pools - A group of objects that will have a one to one (unique) relationship to a service profile.  Examples of this are UUID, MAC, WWPN, and even the physical server itself if you create pools of servers.  Like The Highlander, there can be only one!

Policies - A group of attributes or rules that will have a one to many relationship to the service profiles.  Examples of this would be boot order policy, VLAN and VSAN membership, etc.  Many blades can have this same value.

Templates - A template bundles together various features and attributes into a logical object that can be copied in a consistent way.  Examples of templates can be virtual nics (vNIC's) and hba's (vHBA's) and most importantly service templates.  Templates may also contain both policies and pools.

A mapping of all the objects into a service profile would look something like this:


(For those of you who don't get the Frankenstein, see my previous article on Pools here)

Now, not every server profile needs to be this complex.  Remember, there are three ways to give an object in UCS a value 1. Take the hardware defaults, 2. Manually override the hardware default, or 3. Pull your attributes from a pool/policy/template.

The really sexy part of all of this is the ability to configure all of your typical server deployment features into bundles that can easily be recreated and transferred from one blade to the next.  Need to configure one blade, no problem!  Need to configure 64 blades, no problem!  I have also started to investigate the world of scripting and XML API modification of the service profile objects.  Expect future updates on this subject as well.

Thursday, April 29, 2010

Cisco UCS Pools In Depth

I was playing around with the pools built into Cisco's UCS today and I found some information that I wanted to share.   Let's start at the beginning.   A physical UCS blade is attached to a service profile.  The service profile is the identity of the blade and contains all the attributes of the blade.  My best analogy here is Frankenstein.  The service profile is the magic that makes the physical blade come alive (It's Alive!!).  Until a service profile is associated (think big bolts of lightning zapping the blade), the blade won't power on because it doesn't have an "identity" assigned to it.



You have three options to make this happen.  You can 1. Take the hardware defaults, 2. Manually override the hardware default, or 3. Pull your attributes from a pool.  Pools are in some ways like DHCP.  You could specify an ip address for a server or you could pull one from the DHCP pool. 

What is a UCS Pool?  A UCS pool is a predetermined range for an attribute for a UCS blade.  Examples of this could be the server UUID, MAC address, WWNN, and WWPN's for FC.

When are entries from a pool requested?  Entries from a pool are assigned to a blade at Service Profile creation and live with that piece of hardware for the life of the Service Profile.

How large should I make my pools?  Since service profiles pull entries from pools, you need a pool size to match the size of your service profile environment, NOT the hardware environment.  If you are running stateless UCS, you could have one server with 10 profiles (and 10 entries from the pool for the blade) and pop back and forth between any of them at any time.

When is an entry released? A pool entry is released when a service profile is deleted.  Even if a service profile is unassociated from a server, the entry will remain for future use.

Can you determine the order that entries are pulled? Good Luck with that!  You can but it is pretty complex and seems to pull in HEX (every 16).

What else do I need to know?  The only issue I found is what happens when a service profile is deleted and another is created?  There is chance (although not 100% in my experience) that you may pull an entry that was just deleted.  This could cause issues if end to end clean up wasn't performed on your infrastructure.  Think of a LUN on the SAN that is masked to a WWPN.  If this isn't cleaned up when the service profile is deleted there is a chance you might get yourself into trouble because you could access a LUN that you didn't intend too.  This could be any attribute (WWPN, WWNN, MAC, UUID, etc).

Let me be clear on the last point, this isn't a UCS problem.  It is a procedure issue that exists in any organization.  It is simply amplified with UCS because it is now easier because the attribute is moved from hardware to a pool.  In the case of FC, the problem was at the HBA level and now it is in the pool.

As with everything, with great power comes great responsibility!

Thursday, March 18, 2010

Cisco UCS Just Made VMware VMmark Very Interesting

So, it's time for me to eat a little humble pie...  A little while back I posted how VMware's VMmark has become increasingly less valuable.  This wasn't a knock against the tool, it was a knock against the vendors all using the exact same configurations with some slight tweaks to stay on top.  Cisco UCS in particular took a middle of the road route by not using the things we really care about that make Cisco UCS unique, namely the Extended Memory Technology and the Virtual Network Adapter (Palo card).

Well, that time has come.  Cisco just published a VMware VMmark score based on Intel's new Westmere 6-core processors and.... wait for it...  the Extended Memory Technology and the Palo adapter!!

Here is the link, take a look.  None of the other major vendors have published Westmere processors (that I have seen) so I'm not sure how the scores will stack up.  But, Cisco has done something truly different, because of the nature of Cisco's disruptive technology; they have taken a new road based on their unique technology.  Kudos to them!

Some highlights of the report:
  • Cisco is using the Extended Memory Technology, 192 GB in a 48x4GB configuration
  • Cisco is using the Palo adapter and split one port as one vNIC to the all virtual machine traffic except the web servers
  • The other port on the Palo adapter is split into 27 vNICs.  26 vNICs were presented to the 26 web server vm's.  I assume the 27th was used for the uplink out.
  • I couldn't find anything in the report of where or how the vHBA's were used off the Palo adapter, I wish that would have been included.
Many will say, this is just a "Lab Queen"!! (A Lab Queen is an unrealistic configuration that is dressed up for a benchmark test and would never be used in the real world).  The answer to that question is of course it is!!  But, the fact that this is a DIFFERENT Lab Queen is very cool.  Using a configuration like this is something other vendors won't be able to accomplish and really highlights the ways in which UCS is different.

Monday, March 15, 2010

Cisco UCS QoS vs. HP Flex-10 vNICs in VMware

This post will be more conceptual than technical.  I recently was asked how Cisco's UCS &  HP's Flex-10 network design approaches affect vSphere designs.  Even though the industry is moving towards a unified 10GB fabric, there are different ways to move data through this big "pipe" and still ensure/prioritize delivery.  As you would guess, Cisco and HP approach this problem very differently.  Cisco takes a network centric approach to the problem and HP takes a server centric approach to the problem.

HP's Flex-10

HP Flex-10 takes a 10GB connection and carves it up into multiple virtual NICs.  The size of the "pipes" can be turned up and down to match the amount of bandwidth needed for the NIC.  Think of it as placing smaller pipes in the big 10GB pipe.  This approach is great for vSphere admins because the virtual switches in vSphere can be configured to look just like they did with a bunch of 1GB links into the server.  The transition to this technology is seamless for the vSphere administrator.  I'll borrow a diagram from Barry's awesome article on Flex-10.  If you haven't read it, please do!


What is the down side to this method?
The down side to this approach is by placing multiple pipes within the larger pipes, you have now placed a CEILING on how much data can pass through that particular pipe.  Let's say you present a 1GB vNIC to vMotion and during a vMotion it would be to your advantage to have access to more bandwidth.  Too bad, 1GB is all you will ever get.

 Cisco UCS's QoS

Cisco UCS uses a method known as Quality of Service (QoS).  Most of us "server guys (and gals)" have no idea what this is.  Here is how I have come to understand it.  If this is wrong, please correct me.  Network traffic is given a priority and this priority kicks in WHEN THERE IS CONTENTION on the network.  So, instead of smaller pipes inside a large pipe, you have more of a priority system in place to guarantee certain levels of service.  Think of this as a FLOOR model.  You can have as much as you want as long as everyone else gets their minimums (they get their quality/guarantee of service).  If something needs to spike and there is room, it can spike and then return to normal.  Here is a diagram of our Cisco UCS with traditional switches.  This isn't 1000v but you get the idea.  As you can see, two big 10GB pipes into the virtual switches instead of smaller pipes into multiple virtual switches.


As the vSphere administrator, this looks very different from my old multiple 1GB links into my multiple virtual switches!

What is the down side to this method?

At this time, QoS for Cisco UCS appears complex to configure and represents a shift in thinking for the vSphere administrator. 

How is the QoS implemented for Cisco UCS and VMware?

That is a very good question.  I can't seem to find any documentation on how to actually do this yet.  I'm sure there is a Cisco internal doc somewhere but I haven't found anything public that lays out the hardware that is needed (do I need 1000v or Palo for this, can I use a CNA and the standard switches?) nor have I found a "cook book" that documents how to properly make QoS happen in a vSphere environment.  I'm sure this will happen in time and if you have a link, please leave a comment!

Which is better?

It depends on your point of view and the comfort level of your team.  I can easily see advantages to both approaches.  One is easier to implement, the other appears to be a more elegant (but complex) solution.  Cisco has once again brought a disruptive technology to the table that can't be ignored.  What are your thoughts?

Tuesday, February 23, 2010

Configuring Multiple Palo Adapters in the Cisco UCS B250 Blade

I found out something interesting today that I wanted to share.  Version 1.1.1 of the Cisco UCS Manager code only supports one Palo card on the B250 at this time.  Here is the line from the release notes and the link to the release notes.
  • Cisco UCS M81KR Virtual Interface Card. In a UCS B250 M1 Extended Memory Blade Server, the UCS M81KR Virtual Interface Card must go in slot 0 only and can not be mixed with other adapters. 
This by itself isn't that big of a deal because I can't see many people wanting more than one Palo card at this time.  The problem is that the Netformix configuration tool allows you to configure the B250 with 2xPalo cards with no errors.  I have no idea if you can take the Netformix configuration and turn it into an order at this time (there might be an error check somewhere else in the system that would kick this out) but somehow I doubt it.  Just in case, I wanted to let everyone know.

Monday, February 22, 2010

How Cisco UCS Deals with Split Brains

This will be a short post this morning.  I wanted to pass along how Cisco UCS deals with a split brain scenario.  I'll start by explaining how you would get into a split brain scenario.  In normal operations, one of the 6100's is the active brain and the other is the stand-by.  A split brain in UCS would happen when both of the cluster interconnects betweenthe 6100 Fabric Interconnects are severed (the L1 and L2 ports).  The active brain still thinks he is active and the stand-by no longer sees the active so he tries to take over.  You now have a potential power struggle because both brains think they are in charge.

Luckily the Cisco UCS folks are way ahead of this scenario.  They added logic to the Serial EPROM (SEEPROM) in the UCS chassis to resolve the situation.  The odd number of chassis that are added to a UCS Domain act as judges during split brains.  For example with four chassis, three are acting as judges.  A marker is added to the SEEPROM on these chassis to make them quorum resources.  To clarify this a little further bit more, if there is an odd number of chassis, all of them will be used.  If there is an even number of chassis, it will drop the last one (n-1) so the number of quorum chassis will always be odd.

When the split brain is detected, both 6100's will immediately demote themselves and then claim as many of the quorum resources as possible.  Whoever claims the most quorum chassis wins and promotes himself back to the active manager. The scenario would look something like the following.  Pretty slick!

Sunday, February 7, 2010

Why VMWare's VMmark Scores Have Become Useless

Well, it was bound to happen.  Every time an industry benchmark standard comes out, the manufacturers eventually figure out ways to "cook the books".  I've seen a LOT of FUD flying around from both HP and Cisco lately about VMmark scores and I have been asked a lot of questions about both platforms.  After a taking close look at the scores, I'm ready to throw in the towel.

Before I go further, take a look at the VMmark, 8 cores scores posted here.  You will see the Cisco B200 Blade is on top right now (of the major vendors, I don't count Fujistu, sorry Fujistu) with 25.06 and the HP is next with the BL490 24.54.  A couple of points:

What is the different between really really fast and really really really fast??

What is the difference between 25.06 and 24.54.  Maybe 1-2%?  Honestly, not much if they both meet your needs and you won't be pushing them to their limits.  I'm sorry but that is within a margin of error and/or the test could be reconfigured by everybody to meet the score.  At the end of the day both of them will meet your needs very well and the title of "fastest blade" means nothing!

Both Cisco and HP sell "big memory solutions" but they are no where to be seen!

Take a look at the memory in the details for both of them.  Both the HP 490 and B200 use 96GB of memory.  Where is the B250 with the larger memory footprint? Where is the BL490 with either 144GB or 192GB of memory?  You will also notice that the HP BL490 memory is running at 1333Mhz and the B200 is running at 1066 Mhz.  Since there is no big jump in performance numbers the VMmark score isn't memory bandwidth bound or HP would have had an advantage.  I suspect (although I don't have proof) that the VMmark score is now CPU bound and any memory above 96GB doesn't help the scores.  I further think (again, no proof) that the test isn't pushing the maximum memory bandwidth because there is no change from 1333 Mhz to 1066 Mhz.  It would be interesting to see if the drop to 800Mhz by HP would be noticed in the scores.

Cisco is using an EMC SAN with SSD's on the back end!

Take a look at the EMC Storage section on the Cisco benchmark.  They are using an EMC CX-240 with SSD drives!  There is NOTHING wrong with this, SSD's are coming down in prices but they provide a clear, known advantage to the IOP's numbers that could easily be the sole reason for the 1%-2% increase.  I'm willing to bet that if HP used the same storage configuration, they would produce similar scores.

Why didn't Cisco use the Palo card?

Cisco is using the Q-Logic CNA for the tests.  Why didn't they use the Palo card?  I suspect because it isn't "technically" released yet but that is the benchmark everyone wants to know about.

What am I trying to say here?

What I'm saying is that both HP and Cisco make great products and they will go to great lengths to make the other look bad.  They are so close to each other from a VMmark score perspective that any clear difference can't be shown with the current test.  Don't make a purchase based on a score!

Thursday, February 4, 2010

Cisco UCS - How Many FEX Uplinks Do I Need?

As a Data Center Architect, I have a constant need to know not just HOW to do things, but WHY to do things.  As I dig deeper into the Cisco UCS system I find the concept of FEX (Fabric Extenders) very facsinating.  The number of FEX uplinks may not seem like much, but a couple of cables can have a very significant impact on the design of a UCS system.  If need a refresher on how a UCS system is set up, please see my first article for more information.

What is a Fabric Extended (FEX)?

In very simple terms, the FEX serves as the "pipe" between the blades in a UCS chassis and the 6100 Fabric Interconnects (FI's).  Each FEX has a maximum of four 10GB connections.  Think of it this way, you can "choose" your bandwidth back to the FI's in one, two, or four 10GB increments (the three uplink option isn't supported).  If you plug in one, you get 10GB bandwidth spread over 8 blades.  Need more bandwidth? Plug in the second to get 20GB.  If you really need the maximum then you can go for all four connections for a total of 40GB per FEX.  Remember, each UCS chassis contains two FEX's and each FEX is connected to one (and only one - don't cross connect them!) 6100.  If you plug in all eight connections, you achieve a maximum of 80GB to the chassis or 10GB per blade.  If you are interested in how the traffic flows from the blades to the FEX port (referred to as pinning), here is a great link from Rodney Haywood that details this relationship.  Here is a picture of a FEX close up:
 

Here is how you would cable them (don't cross connect them!!)

Why does this matter?

FEX uplinks directly affect three different areas:
  1. The bandwidth from the 6100's to the chassis
  2. The number of chassis supported per pair of 6100's
  3. The number of vNICs supported by the upcoming UCS Palo blade card.
Let's tackle each of them one by one

1. Maximum Bandwidth per Chassis based on FEX Uplinks -> (FEX's * Uplinks *10)

To calculate the amount of bandwidth available to a chassis: (FEX's * Uplinks *10).  So, If I have two FEX's, each with 2 uplinks, I have 40GB at my disposal (2*2*10).

2. The number of chassis supported per 6100 is inversely proportional to the number of FEX uplinks -> ((total ports - uplinks) / FEX Uplinks per chassis)

Every time you use more than one FEX uplink, you actually reduce the number of chassis you can plug into the system.  Let me use a simple example.

Let's say you have a 20 port Cisco 6120 with the 6x10GB module in it for a total of 26 ports.  To make the math simple, you decide to dedicate six 10GB links for northbound (out of the chassis) traffic.  You can support up to 20 chassis by using a single FEX uplink per chassis.  What if you need a second FEX uplink?  The number of chassis goes down to 10 because you need two per chassis but you only have 20 ports to physically plug into.  If you need need 4 FEX uplinks, then you can only support 5 chassis per 6120.

To calculate the maximum number of chassis, this is the formula to use: (total ports - uplinks) / FEX Uplinks per chassis.  To use the example above, a 6120 with 4 uplinks yields 5 chassis ((26-6) / 4 = 5)

UPDATE:  As pointed out by UnixPlayer, what about uplinks out of the chassis??  Great question and I admit it was late when I wrote this.  I have now updated this to include uplinks.  You of course need some uplinks or this will happen to you!

3. The number of vNICs and vHBAs supported on a Palo card is proportional to the number of FEX uplinks -> ((15*FEX uplinks)-2)

 As pointed out by Kevin on his article, the number of uplinks determines the number of vNICs the Palo card can present.  The thoeretical maximum of the card is 128 but only 58 can be used currently.  This is for both vHBA's and vNICs.

As you can see, there are some interesting design choices to be made based on bandwidth, scalability, and virtual I/O.  I really like the customization ability of the UCS system to tailor to the requirements of the customer but it is also very important to understand the relationships presented above when designing the system.

Tuesday, February 2, 2010

Cisco UCS Information for "Server People"

I've been working with the UCS equipment as time allows for the last few weeks.  I've also had the privilege to visit Cisco TAC for UCS here in Raleigh, NC to pick their brains a bit.  Here is a quick bullet list of some features that I found interesting from my server based perspective.

  • The amount of 10GB over subscription from the UCS chassis to the 6100's (Fabric Interconnects) is proportional to the number of uplinks.  There are four connections maximum per FEX, per chassis.  One uplink will provide an 8:1 ratio, two uplinks a 4:1 ratio, and lastly four uplinks for a 2:1 ratio.  Three uplinks is not supported.  (My next article will be an detailed article on the FEX's)
  • This may seem obvious to the Cisco folks but it wasn't to me.  The 6100's are "backwards".  They are designed to be mounted in the back of the rack so all cabling is towards the rear of the UCS chassis.  Cooling is "front to back" on the 6100's to match the UCS chassis.
  • You can "Mix and Match" adapters cards on each blade because the uplink is a common 10GB fabric.  This means if you only need a few Palo cards in a chassis and maybe CNA's on the rest, you can do that.  Service Profiles won't be compatible but you do have that flexibility
  • Only Cisco memory is supported on UCS blades.  No 3rd Party memory
  • The UCS Chassis needs 2 power supplies.  It ships with zero.  3 power supplies provide N+1 redundancy and four power supplies provides N+N.  As more power supplies are added, the load is distributed evenly across each power supply
  • The UCS Chassis has 8 fans but needs 4 to operate so it is N+N redundant
  • When a Chassis is plugged in, multiple blades are powered up in serial fashion to prevent an in-rush current spike that could blow the circuits.  This has been a problem with other blades customers of mine in the past.
  • The 6100's are active/active for 10GB data but are active/passive for management of the chassis.  At any given time one 6100 is active and constantly passing information over the L1/L2 connections to keep the passive management module up to date
  • The FEX connections on the back of a UCS chassis CAN'T be cross connected to the 6100's.  I'll have more information on this in the next article
  • The UCS Manager allows up to 4 KVM connections at one time.  I'm still checking if this is 4 per UCS Manager, 4 per chassis, or 4 per blade (If you know, please leave a comment!)
  • The maximum number of vNICs the Palo card can present is 56 and is dependent on the number of FEX links from the chassis to the 6100's.  I'm still getting details on this information and I will post this in the near future
  • The 6100 Fabric Interconnects are licensed per port.  The 6120 comes with 8 ports licensed and the 6140 comes with 16 ports licensed.  Additional ports must be purchased individually, kind of like an FC switch.  This applies to both Northbound and Southbound traffic (I REALLY don't like this!!!)
  • Smart Net needs to be purchased on the 6100's, each chassis, the blades, and the expansion modules in the 6100's.  Smart Net lasts for one year so if you want three year coverage, you need to purchase quantity 3 of Smart Net item for each.  This is VERY different from HP and IBM servers.
  • The first three chassis in a UCS domain (managed by the same 6100's) communicate via an SEPROM to verify and prevent a split brain scenario in the event of the 6100's losing communication 
  • The UCS Manager includes the ability to e-mail alerts and all "call home" to Cisco, much like a NetApp storage system

Monday, January 25, 2010

Cisco UCS vs IBM and HP - Where are the Brains?

UPDATE: Thank you to everyone for the great comments!  Please look for the updated sections that I have highlighted below.  I have learned a lot from everyone and I will continue to update this as more information rolls in.  I welcome any and all comments.  Thank you!

As many of you know, my company recently acquired some very nice lab gear for customer demonstrations and proof of concept work.  Many of my peers already know the UCS systems inside and out but I really need hands on to "get it".

As I learn the UCS system I will share my experiences here.  My perspective is to share what is different (good and bad) about UCS compared to the IBM and HP Blade products.  Before anyone asks, I will only be covering IBM and HP.  If you have additional experiences, please share them in the comments.  I also have no intention of picking sides.  At the end of the day I sell and support all of the above systems and I can get the job done with all of them.  They all have their own unique strengths and weaknesses that I intend to highlight.

In case you aren't familiar with what UCS is, I suggest you take a look at Colin's post over on his blog.  He does a great job putting all the pieces together.  Plus, I'm going to steal a few of his graphics. (thanks Colin!!)

A UCS system consists of one or more chassis and a pair of Cisco 6120 switches that provide both the 10GB bandwidth to the blades as well as the management of the system.  The last part of that statement is the key to understanding how UCS is currently different from the competition.  I define management in this example as the control of the blade hardware state.  This includes identification, power on, power off, remote control, remote media, and the virtual I/O assignments for MAC and WWPN's.

By moving the management from the chassis level to the switch level, the solution can now take advantage of a multi-chassis environment.  Here's a simple modification of Colin's diagram to illustrate this point.


(UPDATED!) What are the limitations to the Cisco UCS model?
Someone asked in the comments how this scales.  Honestly that was a great question.  I'm still learning Cisco and I was wrapped up in making it work.  Let's take a look at that.  Currently you can have up to 8 chassis per pair of UCS Managers (Cisco 6100's).  That number will increase in the upcoming weeks and eventually the limit will top out at 40.  But, the more realistic limitation is either 10 or 20 depending on the number of FEX uplinks from the chassis to the 6100's unless you are using double wide blades.  If you don't understand what that means right now, don't sweat it.  I'll be posting about that shortly.


(UPDATED) What if you need to manage more than the chassis limitations today?
If you need to go above the limit, then you have two options.  The first option is to purchase another pair of 6100's to create another UCS System and they will be independent of each other.  The second option is provided by BMC software.  This will allow you to manage more chassis and the solution also provides additional enhancements.  I admit I know little to nothing about the product so I'll just post the link from the comments and you can take a look.  The brain mapping for that would like this.



How do you get into the brains?
Each 6120 has an ip address and both 6120's are linked together to create a clustered ip address.  The clustered ip is the preferred way to access the software.  The clustering is handled over dual 1GB links labeled L1 and L2 on each switch.  They are connected together like this:



Cisco uses a program to manage this environment called creatively enough, Cisco UCS Manager or UCSM.  To access UCSM, point a browser at the clustered ip address.  Once authenticated, you will be prompted to download a 20MB java package (yes it is java, yuck!).  Here is a pic of ours with both chassis powered up.



Notice that both chassis are in the same "pane of glass".  This allows for management of all the blades from one interface and the movement of server profiles (covered later) from one chassis to another within the same management tool.


How does this compare to IBM? 

IBM is a two part answer.

IBM Part One - Single Chassis Interface in AMM

IBM uses a module in each BladeCenter chassis called the Advanced Management Module (AMM).  There can be up to two AMM's in each chassis.  If there are two AMM's, one is active and the other is passive.  They share the configuration and a single ip address on the network.  In the case of failure of the primary, the passive module becomes active and communication resumes on the original ip address.  The AMM will control power state, identification, virtual media and remote control out of the box.  Virtual I/O (both WWPN and MAC) is an additional purchased license in the AMM.  The product is called the Blade Open Fabric Manager (BOFM).  I don't know if BOFM supports 10GB but I know it supports 1GB ethernet and 2/4GB FC.  This is what it would look like with brains in place:


As you can see, each chassis is managed individually.  In my experience, this is the most common configuration I have seen.

IBM Part Two - Multiple Chassis Management with IBM Director

IBM does have a free management product called IBM Director that can pull all this together into a single pane of glass.  The blade administration tasks are built into the interface and virtualized I/O is handled through the Advanced BladeCenter Open Fabric Manager.  Advanced BOFM is a Director plug-in and is a fee based product.  Logically it would look something like this:




The downside to this solution is you now have another server in your environment to manage.  In my experience Director is a little flaky at times but I also haven't tried the newest version which is a redesign to address many of the issues.

How does this compare to HP?


HP is a two part answer as well.  I haven't implemented HP's Virtual Connect over multiple chassis so I will ask that if you know this answer and can throw some links my way, please do and I will update this section.


(UPDATED!) HP Part One - Single Chassis Interface in Onboard Administrator (OA)


HP's approach is very similar to IBM.  HP's management modules are called the Onboard Administrator and there can be a maximum of two in each chassis.  HP is different from IBM because each module requires an ip address.  At any given time one ip address is active and one ip address is passive.  If you access the passive module on the network, it will tell you that you are on the passive module and instruct you connect to the active module.  Like the IBM AMM, the OA will control all basic functions such as power state, identification, virtual media, and remote control.  Like IBM, HP has a separate product for virtual I/O called Virtual Connect.  Unlike the IBM and Cisco products, HP's Virtual Connect is implemented at the I/O module level.  The only way to achieve virtual I/O is to purchase the HP I/O modules.  HP's brain mapping is a little different than IBM because you can connect up to four chassis into one interface.  Since you probably won't be able to power more than four chassis in a rack, think of it as consolidation at the rack level.


(UPDATED!) HP Part Two - Multiple Chassis Interface in HP Insight Tools


After you get to four chassis, HP Insight Tools need to be brought in to fulfill the needs.  Based on the comments below it appears that two products will fit the bill.  To manage the chassis and blade functions you will need Insight Dynamics VSE Server Suite and to manage the virtual I/O you will need the Virtual Connect Enterprise Manager product.  Both the Insight Dynamics VSE Server adn the Virtual Connect Enterprise Manager is fee based.



Summary

(If you made it this far, I'm impressed!)  Cisco's approach feels very "up to date".  I really like the idea of not having to add another server (and additional fees for virtualized I/O) to the environment for management of the products.  By moving all of the management centrally to the switches you are better able to see the environment and implement a multi-chassis/multi-rack solution.  IBM and HP offer a similar solution that has grown over time but the roots of the interface are in single chassis/rack management.  But, at the end of the day both IBM and HP offer a centralized management solution.

Thoughts?  Concerns?  Please leave a comment!

Tuesday, January 19, 2010

That's A Lot of Hardware!

Just a quick post today.  As some saw on Twitter yesterday I will be getting my hands on some pretty impressive hardware.  My company has decided to move our customer demo lab to our office and all the gear arrived yesterday.  Here's a few pictures for now but we will be setting all of this up over the next few weeks.  I will be posting some impressions and tips as I go.  With my HP and IBM Blade background I am hoping to write a good bit on the UCS experience.  In addition to the EMC NS-120, I am hoping to integrate our existing EMC NS-960 for some experience with that hardware as well.  Should be interesting!!

Picture #1 - 2x Cisco UCS Chassis each with 4 blades, 2x Cisco 6120 Nexus, 1x Cisco Nexus 5020


Picture #2 - A LOT of NetApp disk shelves (NetApp controller not pictured)


Picture #3 - EMC NS-120 still in the box (but not for long!)

Thursday, October 22, 2009

Cisco UCS Requires FC To Be Stateless Today

There was a great conversation going on via Twitter yesterday that I wanted to comment on a little further. Sometimes getting things out in the open for "group think" can lead you down some interesting paths. Rod Gabriel was sharing his research regarding UCS and the idea of stateless blades. It all makes sense when you think about it but I just hadn't gone there yet.

In order for UCS to be stateless, you need to "SAN" boot the OS. Local booting of an OS prevents you from a true stateless environment. Since the primary target for UCS appears to be VMware, this means FC is the only way to boot and remain stateless. As Scott Lowe and Paul Richards pointed out, you can PXE boot ESXi but it is still in experimental support. UCS will not iSCSI boot VMware.

As I stated on Twitter, this really disappoints me from a design perspective. The big advantage I see to UCS is the statelessness of the blades but the current requirement for FC to boot ruins much of that value for me. For me personally and my customers, this limits the value of UCS into new clients dramatically.

I haven't installed FC into a new environment in over a year. Most of my new customers choose a server platform local booting and then use an IP based solution to go to storage. In my opinion, UCS will not make much traction into the SMB (Small & Medium Business) market until this changes.

Yes, the market will change and the problem will be solved but until it does, the loss of statelessness decreases a lot of value in my opinion.

I would love to hear your opinion on the subject!