Contact
About Karl Katzke Resume Consulting
Resume
Consulting
Media organizations that are traditionally a day late and a dollar short, like InfoWeek, have been talking again about whether you should keep your organization’s IT infrastructure in the cloud or on local servers .
The analogy that’s used in that article is that many industrial businesses used to generate their own power, and used to keep engineers on staff to support that power generation capability. These days, it’s more efficient for them to tap into the grid, because electrical power companies that can operate at scale will generate power at more economical prices. It’s not a bad analogy.
What was left out of the analogy is all of the other things that came along with that physical plant that benefitted an industrial campus. Most of these plants were steam plants in some form or another; they burned something or used waste heat from somewhere to make steam, which is an incredibly efficient way to transmit massive amounts of motive force over short distances. You can easily tap the line for a steam plant to heat office space and break rooms. Steam plants generally require heat exchangers and cooling towers, which can be rigged to provide chilled water even in the summer, providing a way to cool offices so that executives can wear suits in the summertime.
For most small businesses , the target market for software like Microsoft Small Business Server, in-house infrastructure doesn’t make sense as long as you have a sufficiently advanced or reliable internet connection. (This generally means that you need 2 or more internet connections and an edge device that can handle switching back and forth between them, because even a Metro-E connection provided by Time Warner only comes with an SLA that isn’t worth the paper it’s written on.) If you lose your internet connection during the day because of a fiber cut down the street or because your latency spikes through the roof, you may as well send everyone home.
No, you probably can’t keep running off of the WRT-54g that’s been in the “closet with all the wires” for years and a “business DSL” line from your local telecom monopoly, because they won’t even bother sending a tech out unless your line is dropping more than 25% of it’s traffic.
On top of that, you still need a techie of some sort around to diagnose that and to talk in the language that the telecom monopoly can understand. This function is already filled pretty competently for small to mid sized businesses by numerous Managed Service Partner companies.
For businesses that do software development in-house , the cloud might make sense. Developer machines are typically powerful enough to run Vagrant or another virtualization solution so that developers can test locally. If not, replicating your production environment for your QA process can get moderately expensive if you’ve got a decently sized development team and you do continuous integration internally.
For businesses that do any significant amount of data processing , the cloud probably doesn’t make any sense. You simply require too much compute. Let’s take a look at some numbers. First, instances are expensive . Comparing a compute-optimized node that matches the production environment that I typically specify (m3.xlarge – 4 cores, 7.5 gb of RAM, 80gb of SSD storage) – it’ll run you $0.45 per hour. That’s $10.8/day. Obviously, the reserved rate is much cheaper .
On top of that, as soon as you hit AWS (or DigitalOcean or Azure), the premise that you can get rid of the expensive engineers goes out the window. You need those people, either on a full time or on a contract basis.
I have a consulting client that does a (relatively) small amount of data processing and has their servers in a Tier II colocation facility. For a 1/3 rack, they pay $450/mo for power, 10mbit/s up/down, and cooling. (Yes, that is cheap, they’re grandfathered into a contract rate. The Tier IV side of the datacenter would also be more expensive.) The rack has five 1u machines, a 3u tiered SAN chassis with a flash cache, a 1u chassis with backup drives in it, and a Cisco router and firewall. All five 1u machines are approximately double the specs of the AWS instance I specified and they generally host two or more production VMs and a handful of development or testing instances, which means that the overall CPU utilization is rather high.
For comparison pricing, I was able to throw the production instances into various AWS categories that would meet their needs, but with a big penalty for contract costs, the data in EBS, and the database. I used the EC2 calculator and specified ten m3.xlarge instances at 80% utilization and thirty t1.micro instances that are only on when developers are working (assume 40 hours per week, but it’s probably more like 60 or more by the time you count the continuous integration machines), which is a decent measure of how many dev and staging instances the group has running around. Worse, the data processing instances process data feeds, store them, and regularly re-poll them; this means that they do a lot of IOPS and use a lot of bandwidth during reporting cycles.
Total outlay to purchase the equipment in my client’s rack was $42,000 with warranty, and are guaranteed to work for at least three years, and will probably work for longer. Additionally, at the second year mark, we started swapping in new processing equipment at about $3,000/yr for a total annual cost of $8400/yr. I expect the SAN chassis will need to be replaced at year 7, but that’s outside of the horizon of this discussion. The equipment also has residual value that can be recouped by selling it if it does not fail after we rotate it out.
What gets really expensive is database traffic. There’s a RDBMS that does a big chunk of the work. It requires consistency (eventual consistency is not good enough), chews a lot of bandwidth, and has about a 3:1 read:write ratio. Every pageview in the web app does about 150 queries (yeah, I know, I’m not the programmer) and every piece of data that we ingest gets written to once, read twice, written to once again, and then read monthly from there. Our warehouse is largely flat file at this point.
Total cost for AWS, excluding database, is $11,300 one-time fee for reserved instances, and $2200/mo ($26,400/yr) to run. The real joker, price-wise, in the stack is DynamoDB. By the time you include a fully consistent database, you’re looking at over $10,000 per month. You could buy this client’s existing environment every six months and still save money.
With a thorough re-engineering of the environment to use a more economic data storage method and to minimize database traffic, we could probably bring the cost down to where we’d only be able to buy the existing environment every two years instead of once a year. AWS just doesn’t seem to make sense for this workload. That re-engineering would require a significant investment.
Back on the other hand, I have another client who already uses AWS heavily (SES, others), but primarily uses two leased servers, one for database and one for web. He’s I/O bound during peak time on the leased servers and they consume a significant portion of his hosting budget. Most of the time, though, those leased servers sit idle. The database ticks along at about 10% CPU utilization. The web server gets bound when a traffic spike coincides with scheduled data processing tasks that exchange data with various providers’ APIs and create summary reports.
This environment, which is a traditional webapp, is far easier to engineer for the cloud. We’re standing him up in DigitalOcean. While DigitalOcean is a much newer host with a lower availability level and a less feature-rich toolset (no auto-scaling or beanstalks here), we can do most of the work with $5/mo droplets for the load balancer and a couple of web heads. Their database server suffers from some of the same constraints as the previous case, but we can still run that on a $10/mo droplet. Best yet, we can fire up individual small droplets for the data processing jobs, and only incur a small hourly charge. This is a significant departure from the way the web app is currently architected, but it’s easy to squish all of the pieces into place with a sufficient application of sysadmin glue.
Back on the other hand, that glue gets expensive. I’m not that cheap these days, even if I’m still easy. (To get along with, of course.)
If you understand your environment thoroughly and you can engineer it for cloud computing, you can do cloud computing pretty cheaply.
If you’re working in a traditional office environment with just a file server, you’re good to go for the cloud as long as you are cognizant of what you need to get in touch with that cloud.
My number one fear in my day to day work life as a Systems Administrator who works heavily with data processing workloads is the type of non-technical C-level manager (i.e. a CFO) that would read one of these articles and think that “the cloud” could replace all of the expensive colocation bills, capital expenditures, and salaries that come with doing everything in-house. Non-technical senior management sometimes has the ability to block or stall essential infrastructure purchases. If that C-level had read the InfoWorld article, and I was one of my bosses, I might have a difficult time explaining why our workload isn’t suitable for a cloud environment based on the understanding that the nontechnical C-level has. In the past, I’ve seen that lead to a pretty hostile environment between business management and technology management. Thankfully, that particular discussion is above may pay grade.
Except when I’m working as a consultant. Then I just charge double my rate for that type of work.
Your email address will not be published. Required fields are marked *
Name *
Email *
Website
Comment
You may use these HTML tags and attributes: <a href="" title=""> <abbr title=""> <acronym title=""> <b> <blockquote cite=""> <cite> <code> <del datetime=""> <em> <i> <q cite=""> <strike> <strong>
When Sysadmins Ruled the Earth
Cloud vs. In-House Infrastructure
Back in Blog
Getting Started with F# and Mono on OSX
Hardware Vendors Suck AGAIN
Dave on Thinking about Blades? Downsides to consider…
Cliff Wells on STONITH/Fencing: Why You Need it
karlkatzke on RHEL 5 supports XFS out of box
Enana on RHEL 5 supports XFS out of box
Axel on Best Practices for Managing Releases with Subversion
May 2014
March 2014
December 2012
September 2012
August 2012
April 2012
February 2012
September 2011
May 2011
January 2011
December 2010
November 2010
September 2010
August 2010
July 2010
May 2010
April 2010
March 2010
February 2010
January 2010
November 2009
October 2009
September 2009
August 2009
July 2009
May 2009
April 2009
March 2009
February 2009
January 2009
December 2008
November 2008
October 2008
September 2008
August 2008
July 2008
June 2008
May 2008
April 2008
March 2008
February 2008
January 2008
December 2007
November 2007
October 2007
September 2007
August 2007
July 2007
June 2007
March 2007
October 2006
.Net
apple
centos
fedora
first application
goals
home improvement
howto
linux
meta
mysql
opensuse
packagekit
photography
php
punditry
puppy
python
reading list
renovations
reviews
symfony
sysadmin
Uncategorized
vmware
webdev
wordpress
WPF
zend framework
Log in
Entries RSS
Comments RSS
WordPress.org