Tuesday, May 26, 2009

The Obsolete Blues

After struggling for over a year to get a VMWare solution up and running robustly and then priced out for a recharge service, I find that our central IT group has done the job already and for not much more than we can do it for. Thus, there is really no reason to be in this particular business.

I did see this coming, it is pretty clear that cloud computing is the real next wave. VMWare is what established IT shops do, and will do, to make those 'private' clouds.

The real cloud computing outside of the private world, will be based on higher level abstractions than the Operating System. Since I have been an advocate of abolishing any user interface into an operating system (after all, are we not really trying to perform application logic?) it should come as no surprise that I would also advocate for cloud computing using interfaces or API's well above the operating system.

What people are going to want at the most primitive level is going to be Ruby on Rails servers, PHP servers, Java servers or some other programming abstraction (Hadoop?). Then they are going to want data management, but not file managment, I mean contextually relevant data management. That could be via database systems or it could be via something else. This is the developer level access to clouds.

But what even more people (non-developers) are going to want are applications that just work and do useful things. Contact managment systems, billing systems, mail systems, customer relationship systems, social networking, media sharing and purchasing, etc.....

So, my new career, should I choose to accept it, will most likely be trying to show end users in academic research, how to get what they want out of clouds. I would like to think that I will be part of building a private/public cloud to facilitate the transition, but I don't think that a group of my size can ever effectively be in the infrastructure business. I don't know what I was ever really thinking in trying to do that anyway....vanity maybe.

Friday, May 8, 2009

Everything you know is wrong (again)

More people are finding me on facebook now than twitter. I still don't really know how to effectively use facebook or twitter, so I continue to write this blog since writing lots of words, whether I succeed at communicating or not, is what I seem to be good at.

Yesterday, I signed the petition to encourage my federal representatives to support the use of VISTA as the core of a proposed new bill "Health Information Technology Public Utility Act of 2009"

Many, in the technical communities I travel in, find VISTA's core use of MUMPS as reason enough to ignore it. This is because MUMPS is an 'old' technology and everyone knows that old technology can't be as good as 'new' technology. Or at least we have built a market based on that, with computers 'lasting' just 3 to 5 years before they need to be replaced. That sort of technological imperative thinking is just too simplistic for me anymore.

(Bet you were wondering if I would get back to the title of this post :)

So, what else do we know that's wrong? One thing that really strikes me is the nearly universal notion that 'we have to get this economy back on it's feet'. I take that to mean, at it's simplest, that we need to get back to the way things were! You can see this everywhere: Banks are now making money so those high paid executives who created that innovative engine of growth, financial derivatives (say sub-prime mortgage's), need to be rewarded again. The automotive market simply has to re-structure itself for lower operating costs, as if alternative living, working and transportation arrangements that are demonstrably better along many important dimensions (health, energy consumption) no longer need to be encouraged. Finally, we need to spend a lot more 'stimulus' money to get all those retarded health care practitioners to adopt the latest technology.

Thursday, April 30, 2009

Don't know where to start

Major re-thinks going on regarding the following:

Storage Services
Clouds for the academic research community
Operational costs of IT (OpEx)
Is green the new red?

Some random thoughts:

How come I just found out about the 'Prisoner' remake, it's in post production right now?

No one really understands service levels.

No one ever thinks about risk in a way they can communicate about.

'IT' as a concept space has grown too large to manage. We already knew that as a terminology generator it exceeded the DOD some time ago. But now I find that I hear terms, read articles, read analysts reports and come up with high level internal architectural scaffolds with which I understand what I am talking about and asking for. It's just not likely that the people I am talking to or requesting product use the same scaffold and model. The net result is that I am constantly either disappointed or feeling abused and/or ripped-off.

Thursday, April 9, 2009

The little things strike again.

First, on the "my horse for a drink of water" front, our installation of close to a half million in hardware has been held up for lack of a 1 meter long fiber cable. Now have plenty of fiber cables, LC to LC and SC to SC, but what we don't have is SC to LC and that is what we need. The vendor was supposed to ship us one, but they don't have them either!

If you have not guessed, LC and SC are connectors! SC are bigger than LC and seem to be prevalent in switching gear, at least what we get from Cisco and Procurve. LC are prevalent at the server side of things, such as 10GbE cards or Fiber Channel cards. Yes, you guessed it, the 'enterprise' switch vendors are in an enterprise that consists of switching gear and no servers. How long have servers been needing switches to connect them to the network? As another example of this, just recently ProCurve announced a new switch line designed for the machine room, which mainly means it has front to back cooling, with the front defined as where all the connectors are. Up until now, both CISCO and Procurve switches had either side to back or back to front cooling while all servers have front to back cooling. Try mixing that up in your racks for some turbulent air flow. Of course, this is all a matter of perspective since while the 'front' of switches is the front of the rack for most installations, the 'front' of the server is bereft of any connectors aside from the occasional USB port. Servers put their connectors in the 'back'
So, as a guy who installs servers, I really don't care which perspective is right or 'better', just that everyone agree's!

As a footnote, not all server vendors put their connectors on the back, a rather small vendor (Capricorn) which supplies the Internet Archive with it's servers, put's all the connectors on the front. Seems like they must have visited a machine room......

Now, for something completely different......

The second thing is, you may remember our NFS problems? Look at past postings for more information. You may also remember our 10GbE issues (hardware, drivers) as well. We have now demonstrated on the 1/2 of 10 identical servers which can't transfer much NFS data, that they can't transfer much of any kind of data, FTP, SFTP, etc.....They can, however transfer data just fine over Infiniband and link aggregated 1 GbE links!

So we are certain that this problem is in the 10GbE path. Whether it's software, firmware, or chipset issues, we don't know yet.

Hey, didn't you say these servers were identical? Yes, they all came in on the same shipment, their BIOS and Firmware all seems to be the same and it doesn't matter which OS we load, official Solaris, OpenSolaris, NexentaStor the one's that have a problem have it with all three OS versions.

But, there must be some difference somewhere, more sleuthing is required. By the way, when these kinds of things happen, don't believe the vendors that they are ready to support you, you are on your own.

I know the enterprise switch vendors are trying to tell us that 10GbE is the server room fabric of the future, but personally, Infiniband is far more robust and even has better performance to boot! It's also cheaper! But eventually we need to get to Ethernet to get on the Internet, thanks CISCO and Procurve for ignoring Infiniband and leaving us small fry with endless grief.

Friday, April 3, 2009

LOST - in data storage confusion

Previously on LOST - The data center edition, we were having problems with our particular enterprise implementation of Active Directory and SUN's CIFS server. We are switching vendors in an attempt to keep this project on the rails as the derailment looming ahead would be the "BIG ONE" for us.

To accomplish that we switched our storage node inter-connect protocol from iSCSI to NFS(v3). We, being completely paranoid by now, decided to re-run our stress tests on the X4500 storage nodes using NFS this time. Well, things have gone from bad to worse. So far, out of 8 X4500's, we have had only 2 pass the stress test. The one's that fail, do so anywhere from 200GB to 385GB into the test. The two that work can transfer over a TeraByte with no errors. What we see are client side timeout errors which eventually lock the client side machine, forcing a reboot of it.

(Fast cut to another scene, in another part of the data center)

The Oracle systems folks are starting to scream, their systems are locking up, both client side and server side and nothing, not even a trip to the remote managment ILOM console or a mad dash to the data center to attach a serial cable, will bring these servers back to life. They must be power cycled. What's common here is that the Oracle servers are NFS mounting X4500's.

(Cut to the last scene, a despondent group of IT folks in a darkened conference room, staring at log data.)

The leader say's to everyone - does anyone know the difference between the servers that fail and those few that don't? The answer is not yet. What are our plans? Well, we need to exercise the vendor support path (we have two vendors we can try here and three options). We also decide to take a stab at running the venerable and slow moving official Solaris 10 release on one of our failing X4500's and re-run the tests. We also decide to try and swap in a SUN 10GbE card but we need to buy one and get fiber transceivers and packs for our switches.

The scene fades out with no joy in IT land, things are collapsing faster than we can send in the operatives to repair them.....deadlines are looming, the clock is ticking, we can see the broken tracks in our headlights, .... to be continued.

Monday, March 30, 2009

UCS, Clouds and HPC

How many buzz terms can I squeeze into a blog title? I wanted to add a few more, like ARRA but enough is enough :)

Cisco's positioning of UCS is as a complete data center infrastructure. They say that clouds are driving data centers towards UCS. However, it's not really a complete data center infrastructure because it leaves out most of the facility components. That means it is designed to be dropped in place into existing data centers. I don't think that's where clouds are going, in fact, I think they are going to the companies that do the most vertical integration all the way to the power generation facility.

So, Cisco is trying to sell a radically new IT architecture into established IT shops (those that have data centers fully or partially populated now and are looking for incremental expansion or replacement).

Given the pre-recession resource realities (oil at $150/barrel and heading higher) the incremental improvements in operating expense might have been enough to sway these IT shops. But that world is now gone, at least for awhile, it will come back.

So the cloud vendors might use UCS but probably few will because if you want to survive in the cloud market you need to innovate from top to bottom (think power generation, think building construction) and drive costs as low as they can go.

Clouds and HPC have been the subject of a certain amount of academic debate. Most of the naysayers are those that want to have the last ounce of performance, whatever the cost. As you can tell from my previous paragraph that's not in the clouds...(pun intended). So what is in the clouds is that researchers need to learn how to effectively use cloud resources. In one sense, they are already doing that with grids, like the Teragrid. But in the large scale roll out world, I continually run into researchers and small research teams who follow the 'build it yourself' model of HPC and it's close relative - higher an integrator to build it for you. This is the 100 to 1000 'core' market and while some of it needs the last ounce of performance, the fact is that many of these users can't extract the maximum performance from what they have in the first place. Parallel programing for HPC is just too hard or too obscure. And here is where both the solution and the problem comes for cloud vendors. Showing these 1000's of small research teams how to effectively use cloud resources. How to permanently move large data sets into the cloud infrastructure and thus avoid the nasty performance issues (not to mention billing issues) of multi-Terabyte data set access. How to use functional programming to solve their algorithmic problems, such as MapReduce.

If this sounds more like consulting, then you are right, it is more like consulting and less like buying generic off-shelf X86 servers. And that is also a sign of the issues here, because the paradigm of buying generic, off the shelf, PC inspired X86 servers is well cemented into the bulk of the research community that has not yet moved to HPC.

So far, my experience in trying to set up an HPC support team for Bio-Informatics that is software focused has been met with apathy. Nearly everyone want's to talk hardware and which processor are you using? If I say it really doesn't matter as long as you can get the job done in an acceptable time frame at acceptable costs, well they turn away and go back to browsing the hardware vendor web sites. You just gotta have the latest processor, not using Nehalem yet? well get with it, it's so much better.

I do think these things will change and here is why. Scientific research is a competitve arena as much as building and selling products or services is. If cloud computing can deliver the productivity gains that I believe it can for research computing, then those who adopt cloud computing will soon be out-competing those that don't (of course this assumes that at least some of the world renowned researchers who get big grants sign on to this). It also means, that just like in IT, a small startup (think smaller school here) that is funded well enough to land a top scientist, will probably come to the clouds first. That is because they are not hobbled with existing infrastructure, either in data centers or older HPC systems.

The big guys (think large academic research univerisites) will also 'get it' but claim to not see the demand for cloud computing. I suppose that in the more fully de-centralized universities, less funded researchers might see the opportunity and use it to do research that at one time could have only been done at a few high end places......

Hope springs eternal.

Friday, March 20, 2009

RDBMS, Scientific data and Open Source.

OK, I know I recently said I gave up on open source. I must confess that I meant that comment only in the context of US Health Care software, where powerful entities have too much vested in the current state of affairs. When you have a market that is very large, and Health Care software in the US is a plus $80 Billion industry, then it's going to get a lot of attention. But in research, especially academic research, things are different. Since the current focus of a lot of academic health care research is on 'translational' or bench to bedside, i.e. technology transfer, one must wonder if market forces will stand in the way? But in a blog post on the ACM's website, I found another possible outcome - what will be needed for some of the largest scientific problems of our day, simply will not be developed by the market!

It's been a long time since I let my ACM membership lapse. The press of getting things done with commercial systems sort of made most of CACM irrelevant. But in time, good ideas should make their way from the lab to the product, but in this blog entry by Michael Stonebraker, he thinks that won't happen with scientific data management and furthermore he thinks the problem is too big for any one academic institution - hence an organized open-source project.

http://www.cacm.acm.org/blogs/blog-cacm/22489-dbmss-for-science-applications-a-possible-solution/fulltext

For those who don't know Stonebraker, he is one of the "academic to industry" pioneers in RDBMS land, having been behind the creation of Ingress which later morphed into MS SQLServer.