Showing posts with label Java. Show all posts
Showing posts with label Java. Show all posts

Monday, January 17, 2011

Scala Considered Harmful For Large Projects?

Original Tweet

Last week, I was having an extended email conversation with Geir Magnusson Jr. (yep, that Geir, of Gilt Groupe/MongoDB/Apache Software Foundation fame), where we were talking about whether it was a good, bad, or indifferent idea to allow people to start to include Scala into large-scale programming projects that were otherwise based in Java. Geir and I ended up both agreeing that this was actually a bad thing, and should be avoided for the time being.

I tweeted the above missive, and all of a sudden I had more than 10 people contact me saying "please tell me more!" Let no one think I ignore my faithful readers.

What do I mean by a large-scale programming project? I mean one with the following characteristics:

  • It has a sufficient size that not every developer is potentially even familiar with every module.
  • It has distribution such that in a production use case, it is not a single monolithic process/binary.
  • It has enough people working on it that there is some type of specialization in terms of the teams and/or programmers.
  • It has flux in the developers (new people joining, people being replaced).
  • The same code base has been in constant development for a period of years, or is anticipated to be.

I've written about this in the past, where I referred to my desire for a Journeyman Programming Language. For what it's worth, I'm 100% behind the Backwards Incompatible Java movement spearheaded by one of OpenGamma's developers, Stephen Colebourne. And in that frame, what we're talking about in a large-scale programming project is precisely the context in which a journeyman programming language becomes useful.

So why did Geir and I concur that someone who has a large-scale Java programming project shouldn't start adding Scala into the mix?

Two main reasons:

  • The inherent problems in a split-language project;
  • The sheer flexbility (by design) of the language;

Do You Really Want Two Programming Languages?

Let's say you've got a single programming language in your large-scale project. In that case, you hire people who know that language; all your tooling is specific to that language; you try to make the best use possible of that language.

Any sufficiently complex large-scale programming project will likely have multiple programming languages inherently (usually because you need to use a particular programming language for a particular function or to integrate with someone else's code). But should you try to intermingle them by design?

Once you've done that, you've started to add unnecessary complexity to the project:

  • Do you have to start hiring dual-language programmers, or spend time training them in the language they don't already know?
  • How well does the tool chain integrate? Are there key gaps or differences in operation?
  • Are you going to inadvertently split the team into people who are comfortable with New Language versus people who aren't? Is that going to reduce your staffing flexibility to assign tasks to developers?
  • Is there going to be an impedance mismatch in terms of people or technology in crossing the language boundary? Is that going to cost you?

These are questions that have been faced by programming teams for years, and usually people come down on the side of not allowing programming language proliferation, unless it's part of an effort to upgrade/migrate the whole code base over time from an obsolete language/model to a more modern one. Would you really get that much benefit from Scala to include it at this stage? I doubt it.

Modern Java has a lot of the stuff that's in Scala. It's not as well integrated, much of it is bolted on as an afterthought, and quite a bit of it is based on convention, but you can get a lot of good stuff out of Java written in a very modern style. On balance, I don't think you get that much out of Scala to warrant its introduction into an existing large programming project.

Language Flexibility Bad For Large-Scale Projects

Let's assume that you're working in a programming language like Scala or C++ that lends itself by design to what I would consider to be abuse of the programming facilities (whether it's the ability to write BASIC code inline in the language, the ability to introduce APL-like operator character insanity, or the gratuitous abuse of the compiler that is turing-complete lambda calculus in the templating system). These are, I 100% concur, cool features. It's very very cool to see someone doing something like the BASIC DSL.

The problem is that if you have a language designed to make it easy to do this, you encourage people to do it. And that's when problems start.

Large-scale programming projects need to avoid this type of thing at all costs. In my opinion, there are several key requirements for the codebase for a large-scale programming project:

  • The code must be intuitive and fast to comprehend for a developer familiar with the rest of the project, but not the code in any particular module. In other words, anyone on the team can edit/fix/enhance any other part of code (where they understand the business concepts and/or underlying math).
  • The code must be intuitive and fast to comprehend at any point in the future. In other words, you can't have serious issues looking at old code (that's still in the code base) or old edits (where you're trying to figure out why a change was made by looking at history).

Does an inherently flexible language with kewl DSLs and custom characters and template abuse satisfy these? No. It doesn't satisfy the temporal understanding clause (because a DSL that doesn't get consistently used over time loses understanding), nor does it satisfy the pan-developer clause (because not everybody on the team, even if they know Scala as well as Java, may know the DSL or be able to effectively type in high-order Unicode glyphs).

Ultimately, a programming language that's useful for large-scale programming projects needs to have clear, unambiguous grammar and syntax so that any developer familiar with the project and the language can instantly figure out what's going on. Any features (operator overloading, terse/different method invocation syntax, DSLs) that add time in trying to figure out what a block of code does slow down the project. Sadly, Scala is chock-full of them, and it appears to be considered good Scala style to use as many of them as possible.

Can Scala Work?

Although you may think I'm all about hatin' on Scala, the answer is actually Yes.

Neither Geir nor I consider Scala inherently harmful to a large-scale programming project. It can be a very good addition, if your tool chain supports it and you're willing to train staff so that everybody can handle it.

But you have to have even more rigorous coding standards and review process to make sure that the features of the language that excite the "insanely cool programming techniques" crowd don't get used. You have to do what every C++ programming group has done for years: pick the features of the language and library ecosystem you're going to use, use them in a consistent way, and never ever deviate.

Note that I actually think that if you're kicking off a new project, you owe it to yourself to take a serious look at Scala as a programming language base for that project. But if you already have a large-scale Java-based programming project, I personally think that I would avoid it as much as possible.

By the way, lest you think that a new component or module comprise a new project in my thoughts, you're sadly wrong. A "new project" is one where, in my opinion, the entire technology team is separate from the previous one. In other words, a new codebase altogether. Think: new startup; new Line-of-Business funded application; heads of technology reporting into different parts of an organization.

Wednesday, August 25, 2010

Java Initialization Barrier Pattern With AtomicBoolean

I've found myself starting to use this little mini-pattern. You might find it useful.

private final AtomicBoolean _hasBeenInitialized = new AtomicBoolean(false);
public void expensiveInitialization() {
  if (_hasBeenInitialized.getAndSet(true)) {
    // Someone else has already done the initialization,
    // or is currently doing it.
    return;
  }
  // Do the initialization that I want to be done only once.
}

Much simpler than any of the other mutual exclusion patterns that I've found myself using.

The caveat here is that in the case of multiple threads, it's possible that one thread (that got to the party late) may return to the caller before the initialization is done (if another thread is currently doing the initialization). Therefore, this pattern isn't suitable where the caller has to guarantee that the initialization is done before continuing in a multi-threaded environment.

I primarily use it where there's an initialization method that many parts of the code are going to call as a defensive measure before continuing. Think of it as a simple solution to the initializeIfYouHaventBeenButDoNothingOtherwise problem.

Of course, it well could be that everybody else in the world is already doing single-initialization this way, and I'm too addled from being outside the day-to-day coding world to have caught up.

UPDATE 2010-08-25 : I changed the name of the control AtomicBoolean to make it clearer that this is a multiple-execution-over-time pattern, and not a general purpose synchronization barrier.

Tuesday, August 17, 2010

The Oracle/Google Java/Android suit and a forthcoming blog post

For those of you who aren't aware, I don't just spit out blog posts stream-of-consciousness. I mean, it might seem that way based on my terrible writing style, but I actually work at this stuff.

Many of you will know that I'm a long-time Java developer. I've professionally done a lot of other stuff, but the vast majority of my experience has been in Java. I like the ecosystem, the toolchains, the JVM; I find it a productive environment, and OpenGamma's software is predominantly written in Java.

There's a blog post that I've been working on mentally for months, and in text form for about 2 weeks. I'm planning on publishing it tomorrow. It has nothing to do with the ORCL/GOOG suit. Nothing at all. I've felt the things in the post for ages, well before Sun imploded, well before Oracle bought them.

For various reasons, I'm not going to say what I think about the suit itself. If you want to know, read Charles Nutter's analysis of the suit. Also read Stephen Colebourne's analysis. Both of these will make you smarter. Me? I have nothing to add.

When my post comes out, though, I wanted a vehicle to point all this out and a simple link in case people thought I was talking about the ORCL/GOOG situation. It's messy and complicated and wrapped in ego and profit and law and policy, and Charles and Stephen put the points across better than I could.

Wednesday, December 23, 2009

2009 Predictions Revisited

As promised, I'm coming back to revisit my predictions for 2009 to see how I've done. As predicted by virtue of my fuzziness, I was mostly right on most subjects. Let's go into particulars!

Messaging Breaking Out
I think I did pretty well there. We've still not got a final AMQP 1.0, but just following Twitter and blogs I'm seeing a lot more people, particularly from the non-financial world, starting to use messaging in their applications.

Cloud Becoming Less Buzzy
Complete strike-out here. The same people arguing amongst themselves over what is "cloud" is still going on. That being said, the use of utility computing (as I prefer to call it) is on the rise, and Amazon has come up with so many innovations in the space that it's hard to keep track.

Java Stagnating
Mixed bag on this one I have to say. Java has definitely stagnated, and we still don't have a Java 7. That being said, it looks like the delayed Java 7 may actually give us the chance to see JSR-310 and closures coming into the language, which would be a very positive development.

Java stagnating leads nicely into my next subject (yes, I'm going out of order now):

Non-Traditional Languages Breaking Out
I think I hit this one right on the head. The stagnation of Java, and prominent proponents of systems like Scala and Groovy, are seeing people being more willing than ever to consider these languages the "next Java". Scala in particular has gone from an interesting programming language to one which is seeing mass adoption in enterprises and in web shops.

C# Over-Expanding
To be honest, I have no idea how I did here. I've found myself completely and utterly outside of the C# ecosystem, so I'll have to leave it to one of my faithful readers to fill me in on how I did here.

Social Networking Losing Money
Yep, I failed here. Twitter's probably profitable, Facebook is almost certainly gearing up for an IPO. I was completely wrong here.

But it's not just the social networks themselves, social gaming has gone from an interesting idea to one that makes lots of money, indicating that the space of social networking has started to turn profitable not just for network providers, but also for network ecosystem partners.

Sun Radically Restructuring
I think I can say I was right here, in that they're radically restructuring themselves into Snoracle.

Next year's predictions on the way!

Monday, October 26, 2009

Java APIs and Ivy: Click-Through Agreements Suck

We've been using Ivy for project dependency management at OpenGamma since day-one. It was one of the earliest decisions we made, and I'm quite happy that we made it. Furthermore, the work that the Ivyroundup guys have done to package up Ivy-specific files has been fantastic.

That being said, there's one big, honkin problem with it: Sun and their click-through agreements to download almost any of the "interesting" APIs that you need for enterprise software development. These packages (like the JMS or JavaMail APIs) are necessary for pretty much anything that you might want to do on a real system, and you can't legally put them up for download. And unfortunately, there are a lot of them.

The workaround, of course, is that someone in your organization downloads them all, clicks-through on the agreement that is completely and utterly unenforceable, and then sticks them on a private Ivy repository and you configure your ivysettings.xml file to refer to that location, but that's not particularly scalable. And for open source projects, it just doesn't work at all.

Yeah, we've got this Ivy thing so that it automatically identifies all the dependencies and downloads them so you don't have to do anything but have Ant installed. Except for the Java stuff, which you have to manually download and put into these particular locations, and that's a 20 minute task minimum.

Kinda takes the shine off the whole experience, doesn't it?

Now given that everybody seems to know that you can handle licensing restrictions in different ways (by a disclaimer in each file and a copy of the license included), why does Sun still live in the past? Why do they require a clickthrough for something that is essentially just an explicit disclaimer of warranties and liability? Why can I download Eclipse or NetBeans without a clickthrough but not the JMS 1.1 API interface files?

Man, I hope the EU finally lets Oracle buy Sun and put this type of stupidity out of its misery.

Tuesday, August 25, 2009

The Ivy Tutorial Strikes Back

We're continuing on with the Jim Moores Guest Post series, as Ivy proceeds to destroy his source code directory.

After my previous problem with the build.xml file supplied in the Ivy tutorial I hit another issue that I thought I'd share. The default ant build file given in the Ivy tutorial generates an example source file within the ant file and then builds it. I, and I presume many others, based their new ivy-aware projects on that file. It has the default targets for downloading and installing ivy. Yeah. Problem is that one of the targets looks like this:

Can you see where this is going? Note where it says to delete the entire source directory. Yeah. That just happened to me.

I lost about half a week's work. It was my own fault. Because I'm a Perforce refugee and we've decided to try using git, I was more hesitant than usual about checking non-working code into my own code branch just because i wasn't so familiar with the git style of doing things. So while I was trying to get the ivy problems ironed out I ran ant clean which promptly deleted all my source code. All of it. Well, except the unit tests.

Please, let this be a lesson that I've learned so that you don't have to. While I accept that this was my fault, I can't help thinking that as a rule no build file should ever delete the entire contents from the src directory, particularly not one shipped as part of a tutorial that one would attempt on existing code.

Ant, Ivy 2.0.0, and java.net.UnknownHostException

This is a Very Special Guest Post by Jim Moores, another one of the OpenGamma team. We thought this would be useful to the interwebs, and until we have the OpenGamma blog ready to go, we're all ranting as one.

Ivy is a great way of avoiding shipping all the dependent libraries with your code and making upgrading dependencies smoother and easier, but it can still be a bit rough around the edges. For those of you new to Ivy, it's a dependency manager that is built on top of Ant that goes off and downloads and installs all your dependent libraries given a pretty simple list in an xml file. It hit version 2.0 a few months back and started to be really useful with more and more projects switching over to it. Of course Maven has done many of the same things for much longer, but Ivy doesn't require the whole head-warping change of viewpoint that Maven does. To be fair to Maven, I've only used it in passing when building the MyFaces JSF source code and while I did get it working, and was impressed, it took a bit of fiddling. My general understanding of Maven is that it's more declarative than 'normal' build systems - you tell it what you need and you have to rely on it figuring the whole thing out and it requires you let go of the idea of targets and explicit flow control (apologies to Maven fans if any of this is inaccurate). Ivy requires less of a leap of faith because it fits in with the Ant style that is more familiar to many but still brings the advantages of automatic dependency tracking. It can even install itself.

When I first tried Ivy three or four months ago I quickly found that some of the packages I wanted weren't available in the default ibiblio maven repository, in my case it was BerkeleyDB. I soon came across the excellent Ivy RoundUp repository. Ivy RoundUp doesn't actually host the packages themselves, it's a meta-repository that hosts information on how to download extract and repackage artifacts on demand. Unfortunately I couldn't just switch over to RoundUp because I needed some packages that were only on ibibio, so I had to delve deeper into the ivysettings.xml file to set up what is called a chain resolver. Chain resolvers do just what they sound like they do - they try one repository and then fall back to another if the package can't be found in the first. Here is my settings file:

Unfortunately this didn't work at all. It turned out that the example ivy build.xml in the tutorial contains the line:

  <property name="ivy.install.version" value="2.0.0-beta1">
</property>
And that 2.0.0-beta1 has a bug in it that means that chain resolvers don't work. Once I updated the ivy.install.version property to the latest version (2.0.0 at the time) it worked fine. This is still a problem with the tutorials, so be aware.

Since then I've needed to start using Hibernate in my application so I thought adding it would be easy as it's one of the most widely used Java packages. Err, no. After finding Hibernate in RoundUp easily enough (the latest ibibio version is really old), my build hung for some time, apparently unpacking a file, and then went totally exception-tastic on me.

[ivy:cachepath]
[ivy:cachepath] download.1.N65558:
[ivy:cachepath]
[ivy:cachepath]       [get] Getting: file://C:/DOCUME~1/Jim/LOCALS~1/Temp//jta-1_1-classes.zip
[ivy:cachepath]
[ivy:cachepath]       [get] To: C:\Documents and Settings\Jim\.ivy2\packager\cache\jta-1_1-classes.zip
[ivy:cachepath]
[ivy:cachepath]       [get] Error getting file://C:/DOCUME~1/Jim/LOCALS~1/Temp//jta-1_1-classes.zip to C:\Documents and Settings\Jim\.ivy2\packager\cache\jta-1_1-classes.zip
[ivy:cachepath]
[ivy:cachepath] C:\Documents and Settings\Jim\.ivy2\packager\build\javax.transaction\jta\1.1\build.xml:53: The following error occurred while executing this line:
[ivy:cachepath] C:\Documents and Settings\Jim\.ivy2\packager\build\javax.transaction\jta\1.1\packager-output.xml:28: java.net.UnknownHostException: C
[ivy:cachepath]         at org.apache.tools.ant.ProjectHelper.addLocationToBuildException(ProjectHelper.java:541)
[ivy:cachepath]         at org.apache.tools.ant.taskdefs.Ant.execute(Ant.java:418)
[ivy:cachepath]         at org.apache.tools.ant.UnknownElement.execute(UnknownElement.java:288)
[ivy:cachepath]         at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
[ivy:cachepath]         at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)
[ivy:cachepath]         at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)
and another ten pages of related stuff which I'll spare you. There was clearly some problem in downloading the JTA 1.1 classes, which are a dependency of Hibernate. I tried upgrading to the latest version of ivy - didn't work. I then assumed it was an ivy package file and went digging around. When I looked to see the temporary file it was trying to open it wasn't actually there. Hmm. Just before that it implies that it's downloading the file in question. Could it be deleting it? I eventually decided it must be a problem with the file in the repository, so I went looking for help there. That's when I came across this thread on the roundup mailing list archives, and then onto the Roundup discussion of manually downloaded software.

It seems some of the software DOES NOT get downloaded automatically by ivy. Unfortunately, instead of giving you a nice error message telling you where to put it, you get a java.net.UnknownHostException. It turns out you need to download the artifacts manually from the Sun website and put them in the temporary directory ivy uses to store package downloads.

Once it did this it all works magically. Apparently the reason for this is that the download requires ticking a license agreement box which can't be done automatically. I hope this knowledge helps someone save some time. I think in the long run it'll probably mean me having a private repository for those artifacts, but for now I'll leave that for another day.

Monday, April 27, 2009

Sybase JDBC Drivers Ignore Your Database Selection

A very common strategy in doing database-driven functional tests is for each user to have their own database instance on a shared test machine. This means that they can do whatever they want on that database instance without impacting other users, and that multiple people can run potentially destructive test cases simultaneously.

A common approach to handling that is to append the user name to the database name, and use a system property to choose the correct database instance for that test run (for example, naming your DataSource bean in your Spring context.xml something like db_${user.name}). You then have the JDBC connection declared to have the user name in the connection string (a la jdbc:foo:bar:blahblahblah/myapp_${user.name}). All works great.

Except with Sybase.

In its voyage of annoying developers who have to use their managed language drivers (the C#/ADO.Net ones are particularly noxious), Sybase have decided to screw you by ignoring this parameter when it's not valid.

The Sybase low-level protocol for this scenario largely consists of the following steps (completely simplified):

  1. Connect as user Foo
  2. Connection is now bound to the default database defined for user Foo
  3. Issue a use correct_database statement
  4. Connection is now bound to the database specified in step #3

Here's the problem: if step 3 fails because the database desired doesn't exist, Sybase will silently swallow the failure, and you'll end up in the default database for the user, and not tell you (seriously, there's no log or exception or absolutely any sign of what database you're in). Even better, the Sybase driver doesn't even store this in any field that you can introspect in the debugger, so the Sybase client library thinks it's in the database instance you want, when in fact it's not.

If you have a shared test database (with some real data in it for doing manual/gui-directed testing), and surrogate databases for each user's "blow-away-the-world" style testing, this is Very Bad Behavior Indeed.

Solution: Stop using Sybase. Seriously. Yes, it may have been en vogue 10 years ago, but it sucks.

Tuesday, April 21, 2009

Oracle + Sun : A Java Perspective

Smarter people than I have written about this, but having worked at several of the major players here (a summer at Oracle proper, 12 months at BEA working on WebLogic Server, 2 years working at M7, which got bought by BEA, which got bought by Oracle), I figured I'd pipe in my $0.02.

My #1 concern here is that Larry is going to attempt to use everything in the IP arsenal that he's just acquired to screw IBM. He's done it before, he'll probably continue to do it. Considering the amount of investment that IBM has made into the Java ecosystem (at least as much as BEA, particularly when you consider Eclipse), and the amount of hostility between Oracle and IBM, this wouldn't surprise me one whit. Towards, that end, here's what I'd like to see clarification on:

  • Eclipse/SWT. For too long Sun's ridiculous love-fest with NetBeans and Swing has blocked any reasonable approach towards dealing with Eclipse and SWT. The worry that I have here is that since Eclipse == IBM in many people's minds, and both JDeveloper and NetBeans are now under the same company umbrella, Oracle may decide to let commercial considerations (sticking it to IBM) and staff considerations (keeping Sun developers who have a thing for NetBeans, plus everybody internally who's committed to JDeveloper) override ecosystem considerations (we all like Eclipse way more than NetBeans or JDeveloper). My request: please properly support SWT. Not to the exclusion of Swing, but don't fight against it.
  • OSGi. OSGi is probably tainted by Sun as being an Eclipse-technology, and therefore hurting NetBeans and helping IBM or some similar ridiculousness. Support it, please, and kill off anything whose sole raison d'etre is to replace it with something lamer. Modularizing the JRE is one thing, providing something completely useless except to replace OSGi is something altogether different.
  • Open Java Implementations. Stephen Colebourne has been talking about this at length, and blogged about the handover. Please don't be such jerks on the JCP.
  • Open Source Projects. Whither Glassfish and Metro and all the rest of it? They compete with existing BEA/Oracle assets, but are extremely valuable in making the ecosystem valuable enough to allow Oracle to extract value from their proprietary assets.
  • API Neutrality. One thing Sun has been good at, because they've never had any world-beating middleware or application infrastructure technologies, is helping to craft APIs that are by and large vendor neutral (such as JDBC and JMS). Oracle, however, doesn't. Is Oracle going to follow a path like Microsoft has with the .Net APIs, where there's enhanced support for whatever Microsoft is shilling and second-class support for everything else (ADO.Net anyone?), or is it going to realize that supporting the ecosystem means vendor neutrality as much as possible? Sun had no choice as all their app infrastructure is second-rate at best, but Oracle has a choice.
  • The Whole JCP Itself. What's going to happen to it? How is Oracle going to behave in general? Will we see projects falling out from under the JCP umbrella going forward?
  • ZFS [1]. ZFS rocks. Massively. But Oracle's been working on Btrfs in part because Sun refuses to allow ZFS to be licensed in such a way that it can be included in the Linux kernel. Given Oracle's investment in Linux, we can has ZFS in Linux?

Mostly, I think we still have yet to see whether Oracle is going to behave like the BEA side, or the Oracle side: is Oracle going to help build the ecosystem, or is it going to use its new IP assets for proprietary advantage against IBM and Microsoft? I hope it's the former, but only time will tell.

Footnotes

[1]: Yes, it's not Java. But really, I can has? Plz?

Thursday, January 15, 2009

Qt Going LGPL: The Java Angle

(One of Many articles on Qt allowing LGPL licensing). Assuming that this includes the Jambi components (Java/Qt bindings), which I can't find any reference to positively or negatively, but since it's part of the language bindings and not the value-add tools I can't imagine they wouldn't, I think this is a great move for the Java community in terms of rich GUIs.

No, I'm not going to talk JavaFX. Ignore that as a distraction at the moment.

Right now, if you want to roll a thick/rich client (and yes, there are still a lot of applications that are (shock! horror!) not delivered through the browser, and that's going to stay that way) in Java, you've got a few choices:
  • Swing. Gag. Even with Nimbus (only available if you've got 6u10 pushed to the desktop), it's ugly, and no matter how good the Windows L&F gets, it's never going to look native in any way. Swing is also an example of precisely why Java needs closures.
  • SWT. SWT would be much better if JFaces was more advanced (I don't want my apps to look like ass, but I don't want to go back to AWT-level programming either), and SWT was more documented, and there were really people using it for things other than Eclipse plugins. Yes, I know there are some other uses, but not a ton to be fair, and so not a lot of example code to copy.
  • Qt/Jambi. Very well done MVC programming model (signals and slots, but with no preprocessor, so it's all Java toolchain friendly), looks native. Oh, and you can embed WebKit.
Honestly, I think this is a really positive thing for Java to have a rich GUI framework that's easy to work with, and looks good. Lord knows we've been waiting long enough for it....

Thursday, January 08, 2009

RabbitMQ/AMQP OSGi Integration

I had a chance to meet up with Neil Bartlett today, and got to talk to him about his RabbitMQ OSGi integation (note that I believe that everything he's done thus far is actually applicable to anyone using an AMQP broker that is compatible with the RabbitMQ Java client, so I'm treating it as a generic AMQP/OSGi module). While it's a work in progress, it's incredibly cool. It allows you to:
  • Declare production/consumption parameters (exchanges, queues, hosts, credentials) in a declarative way;
  • Have your OSGi container connect the dots so that your declarative bundles automatically match up with concrete connections to your AMQP broker of choice;
  • (If you're using Equinox) Use the OSGi shell to interact with the RabbitMQ/AMQP client library.
The first two parts were the types of things that I had worked on in the past (having a declarative way to specify messaging endpoints, down to the broker connectivity level) so that you could change them at execution-time. The third was the real WOW-worthy point.

Assume that you're coding against an OSGi runtime. Assume that you're working with bundles, and dynamically loading and unloading them as you're compiling your code (meaning that you're probably at an osgi> prompt). Assume that want to see very quickly whether your code is interacting properly. Now you can easily setup your bindings and publish and consume messages as your real code is in flight, and see the effects in your debugger.

How amazing is that? Imagine the human cycles that you'll save being able to do all this from one area, that you're already familiar with?!?!

This is a work in progress, and hopefully Ben Hood and I had a chance to give him some places to go from here, but I think this is exactly what the promise of AMQP offers: the chance to do amazing things based on not having to support a million different client libraries, but to code against a standard and make great stuff happen.

If you're using the RabbitMQ Java Client and any type of OSGi container, I highly encourage you to take a look at this and get involved. I see great things happening here.

Wednesday, December 31, 2008

2009 Predictions

Just a quick set of predictions for 2009. I'll revisit nearer the end of the year to see how I've done.

Messaging Breaks Out
Right now asynchronous systems, particularly ones using message oriented middleware, are still pretty fringe. However, even though I disagree with the technology choice, techniques like XMPP are starting to bring asynchronous communications to the masses. I predict that the extension of XMPP, the forthcoming standardization of AMQP, new infrastructure being developed (like RabbitMQ, OpenAMQ, and Qpid), and extension to new programming languages and application execution environments will all drive more and more developers to finally end polling for updates.

Cloud Will Become Less Buzzy
Right now Cloud/Utility computing is really too buzzy to make heads or tails of, and I think it's suffering from that: when you say "Cloud Computing", it means both nothing and anything. I think 2009 will be the year that people start to solidify what works best in a utility computing space versus a local hosted space, and where hosting providers get mature enough and technology becomes mainstream enough that there's an assumption of utility over local hosting.

Java Will Stagnate
I predict that Java 7, if it even hits by the end of the year, will be pretty anemic and, quite frankly, lame. There will be some interesting enhancements to the JVM, but that's pretty much all that's going to be noteworthy.

C# Will Over-Expand
The entire Microsoft ecosystem will expand and grow beyond the ability of typical Microsoft stack developers to keep up. We're already seeing it with WinForms vs. WPF, LINQ vs. ADO.NET Entity Framework, and whole new technologies like WCF. I predict by the end of the year the Microsoft ecosystem has become so bewildering that while leaders in the community push for more change, day-to-day developers push for a halt to just absorb what's been happening.

Non-Traditional VM Languages Will Break Out
Right now most developers in enterprise systems are coding against the "core" languages (C#, VB.Net, Java) against their respective VMs. This reflects in no small part the maturity and power of their underlying runtimes (the CLR and JVM respectively), but the leading edge of the communities have already started to explore other languages that offer concrete advantages (F#, Scala, Groovy) on the tight runtimes that are already available. This will be the year that those actually break out of the leading edge and into mainstream use.

No-One In Social Networking Makes Money
I still think, by the end of 2009, there won't be a player in the social networking space that is cashflow positive, and I don't think it'll be by choice. I think the technology is so disruptive that we still haven't seen the "right way" to monetize it yet.

Sun Radically Restructures
Sun can't keep burning through cash the way they have been, and they can't continue to have such a chaotic story. At some point in 2009, Sun will relatively radically restructure itself in a bid for survival. Hopefully they'll have read my analysis (part 1, part 2). At the very least, they'll change their ticker away from JAVA.

I think you'll note that all my predictions are fuzzy. That's so that in December, 2009 when we revisit them, we can determine that I was partially right on nearly all of them.

Monday, November 10, 2008

IEEE 754 Floating Point Binary Representations

Just to gather up a whole bunch of stuff I had to slog through and make this more googleable, allow me to summarize some various trivia having to do with bitwise representation of IEEE 754 floating point values across platforms. This is primarily useful if you need to read and write floating point values from byte arrays or binary network streams across platforms, particularly if you have to interact with Steve Ballmer's Insanity.

First, there is no official standard for endian-ness when transmitting IEEE floating point data over the wire. That means that Java ends up defaulting to in DataInputStream and DataOutputStream to big-endian format (to match the fact that everything is big-endian), C# defaults to host-endian format (always little-endian in practice, as the Mono guys have learned.) for BinaryReader and BinaryWriter. First point of fun.

Secondly, IEEE 754 floating point representation defines an entire range of values to represent NaN, not a single value. Java takes the approach to make things byte compatible in the wire format by always emitting a single constant value for all NaN values (where all the meaningless bits are set to 0), while C# allows whatever cruft happens to be in the value on the CPU to flow through to your binary representation. And don't assume in C# that double.NaN has all those set to 0. It doesn't. In practice, double.NaN in C# is full of cruft.

This is fine if you read in the value and call IsNaN on it, but not so great if you want to check that your serialized/deserialized byte arrays are fine. For that, you need to mask out to ensure that you're always writing a canonical representation of your NaN values.

A useful C# block if you find yourself having to deal with this stuff is the following (using this will ensure that your binary representations are always bit-equivalent with the Java formats):


Friday, November 07, 2008

Linux Fork Performance Redux: Large Pages

After a comment from Kostas on my last Linux Fork Test post, I worked with my esteemed colleagues to try things out with large page support. Wow, what a difference that made on Linux.

The theory here is that forking involves playing around with TLBs entries quite a bit. Since a monstrous full heap will have a lot of 8KB page TLB entries to contend with, if we shrink the number of TLB entries by a factor of 256 (by working with 2048KB pages rather than 8KB ones), you'll limit the amount of time that the kernel is spending mucking with them.

First of all, make sure you read this: The large memory support page from Sun. Now that we've gotten the formals out of the way, here's some fun we had.

First of all, a stock RHEL 5.2 installation has a HugePages_Total set to 0 (cat /proc/meminfo). No huge pages whatsoever. So you need to bump that up. For my test (maximum 2GB fully populated heap on a 4GB physical RAM system), we decided to set that to 3 GB, which is 1536 2MB pages.

echo 1536 > /proc/sys/vm/nr_hugepages isn't guaranteed to actually do that, and the first time we ran it, we ended up with a whopping 2 HugePages_Total. Second time bumped us up to 4. So we went on a process hunt to eliminate any processes that were stopping us working, and got things down pretty small. Now we were able to get up to 870, which was good enough for my 1GB tests (which indicated the major performance degradation anyway), though not for the 2GB test. (Yes, I know that you're supposed to do this on startup, but I didn't have that option so we did what we could).

And so I kicked things off with the -XX:+UseLargePages flag. Fail.

Every single time I got a Java HotSpot(TM) 64-Bit Server VM warning: Failed to reserve shared memory (ero = 12). And nothing would run. Well, damn!

Turns out those little tiny bits that they say in the support page about not working for non-privileged users are completely accurate. These all went away when I had someone with sudo rights run the process as root, and all my numbers are from running as root. So just assume that even 1.6.0_10 ain't going to allow you to allocate any large pages if you're not root.

So he ran things as root (and I re-ran things as non-root without the UseLargePages flag since I changed the test slightly). Here's some fun comparison:






Heap SizeLarge Pages TimeNormal Pages TimeSpeedup
128MB3.589sec13.217sec3.68 times faster
256MB3.62sec15.314sec4.23 times faster
512MB4.638sec36.692sec7.91 times faster
1024MB3.885sec67.062sec17.26 times faster

Oh, and that speedup between 512MB and 1024MB? Completely reproducible. Not sure what precisely was going on there, I'm going to assume my test case is flawed somehow.

It's happening so quickly at this point that I'm quite suspicious that all I'm measuring is the /bin/false process startup and teardown performance, as well as the concurrency inside Java. I don't actually think I'm testing anything of any meaningful precision anymore. Maybe at a few million forks or with higher concurrency, but I've achieved essentially a constant amount of time spent forking, so I've gotten out of the heap issue really.

So it turns out that you really can make Java fork like crazy on Linux, as long as you're willing to run as root. And I don't know why and my naive googling didn't help. If someone can let me know, I'd really greatly appreciate it.

Did any of this help Solaris x86? Not one whit. Adding the -XX:UseLargePages flag (even though Solaris 10 doesn't require any type of configuration to make it work) didn't improve performance at all, and Solaris was still twice as slow as Linux without the flag.

Friday, October 31, 2008

JAXB 2.x and @XmlElement(required=true)

We've been quite happily using the JAXB reference implementations for a while now, until someone actually evaluated whether the XML that it's generating is valid. Turns out that it's only sometimes valid with respect to your schema.

(FTR, this is going to be a bit tricky since I'm using Blogger and thus don't have the best inline XML-and-code support). (Version notation: This is all valid for both 2.0 and 2.1 versions [up through 2.1.8], but is not valid for 1.x JAXB).

The Problem: Required Strings
Assume you have an XSD complexType which contains two elements:
<xs:element name="foo" type="xs:double"/>
<xs:element name="bar" type="xs:string"/>

In your XJC-generated Java code, both the Foo and Bar properties will be annotated with an @XmlElement(required=true) annotation. In addition, Foo will be declared as a double rather than a Double, so you'll always have a value of some form in the generated class. The problem is with bar.

bar will be declared without any default value whatsoever (unless you're using the JAXB Default Value plugin). Even worse, if you run through a marshall/unmarshall pass on objects or XML that lack a bar element at all, it works entirely fine (even with the default value plugin) and you'll be just fine consuming data without a bar element whatsoever. So you can quite merrily generate XML which doesn't adhere to your XSD, and consume XML which also doesn't. (for the bar element, you'll always have the default value of 0.0 if the XML you're unmarshalling lacks the element).

In general, I'd consider code which is doing this to be bug-ridden, and in need of fixing: if your application knows that you need a bar, why aren't you adding one to your objects before marshalling them? Conversely, you should always try to make do with whatever crap someone barfs at you over the wire if you can (Postel's Law and all that). But it's still confusing behavior.

Workaround One: Make It An Attribute
If you have control over your schema, making it an XML attribute will solve this problem immediately. They're handled differently by JAXB, so it always works as you'd expect.

Workaround Two: Explicitly Validate
At runtime, for performance reasons, JAXB by default doesn't actually have access to your XSD. All it works with are the (generated|hand-written) Java files with annotations (this is a difference to the way the old XMLBeans worked, where it would pre-process a binary form of your XSD for runtime validation). You can make JAXB perform strict validation, but for performance reasons it won't by default.

The problem is that it means that you have to ship your XSD files with your application and use SchemaFactory out of the javax.xml.validation package to load up the XSD into an in-memory Schema instance, and pass that to your Marshaller and Unmarshaller's setSchema methods. It'll slow you down, but you'll be guaranteed to be valid.

A reasonable option at least to me would be to validate in debug cycles (using the assertions system) and turn it off when you know your applications actually all work together happily.

Reader Note: This was written largely because googling this topic never came up with any actual explaination. Therefore, I wanted this to be keyword google friendly. And since I was going to have to document this for other people at my company, I figured I'd document it for the world.

Tuesday, October 21, 2008

Java5 vs. Java6: ThreadPoolExecutor Difference

If you do thread-pooling Java without calling into Executors.new* (e.g. you're constructing your own instances of ThreadPoolExecutor) and you're using LinkedBlockingQueue so that you can build up a task list, be aware that ThreadPoolExecutor got completely rewritten in Java 6.

Under Java 5, the behavior is that if you set the core size of the thread pool to 0 threads, then the pool won't actually kick off a single thread until the offer method on your BlockingQueue refuses a new element. In the case of LinkedBlockingQueue, that means that it won't kick off a thread until you've reached the maximum number of elements in your queue (the number in the constructor that you are obviously specifying to a reasonable number for your workload), which isn't quite what you would expect.

The workaround is to set your core size to >= 1, or to just use Java 6, which doesn't have this behavior (the execute(Runnable) method was rewritten and has special logic if the pool size is 0 to force it to create a new Thread in that case). In general, though, setting the core size to at least 1 is always a safe thing to do and will work across both Java versions.

Yes, this did bite me. And yes, thanks to continuous integration checking JVMs that my coworkers insist on using but which I abandoned years ago, I tracked this down to precisely this problem.

Updated 2009-06-25: Got corrected in a comment that I said SynchronousQueue when I should have said BlockingQueue.

Friday, August 22, 2008

Tiny Java Containers: None Of Them Do What I Want

Dear Lazyweb,

Although I find it difficult to focus on programming while BBC is live broadcasting Olympic Table Tennis, focus I must. But I have a conundrum: I want a Java code container, but I don't like any of the current methodologies that I'm familiar with. Please help.

Yes, I've looked into Spring and Pico/NanoContainer and all that, but their focus appears to entirely be below the level of what I'm looking at. Their focus appears to be running inside some type of application, rather than launching it (so you configure launching Spring within your Web Application: you don't start Spring at the command line). Some of my problems have to do with packaging and deployment of multi-jar projects, and the IoC containers don't work at that level that I can see.

Plus, while I totally buy IoC as a general principle and use it on a regular basis, I don't need a generic configuration file to do it, particularly not in test cases. If I want a configuration file for a particular context, I bloody write one. It exactly represents my configuration options, is typesafe, has a schema, and can be understood by not-me for later configuration. A configuration file that is capable of representing the entire universe of Java types and construction options is essentially programming in XML. For people who seem to hate J2EE, the IoC-container-crowd seem to have adopted its worst tenet: replacing simple, clear, short snippets of source code with a monstrous amount of generic XML.

But I digress.

Anyway, here's what I want:
  • Packages up a set of jar files as a single classloading context. This is pretty important because I think a big problem with modern composed Java applications is assembling all your various jar files and dealing with diamond dependencies. But anyway, I want something that I can pass around and say "this is the code for my application."
    Yes, I know OSGi is a better solution for this, but I'm not there yet, and I won't necessarily be there for quite some time in terms of OSGi-ifying everything I've got. Bundle of JAR files I can handle.
  • I want to be able to load multiple of those jar bundles at any given time in the same JVM.
  • I want to be able to run static methods (maybe even main itself) with arguments. Heck, it's my code, I'll even adhere to some type of contract (whether explicit interface or implicit through annotations) to give me better lifecycle support.
  • I want to be able to run lots of instances of those static methods (think message processors, each instance on a different JMS Topic). I want that to be dynamic (so I have one service kicking off others).
  • I want to scale those instances on multiple machines. I'd like that to be automatic.
  • I want to automatically handle failures of machines by distributing the load on that machine to other machines.
  • I want lots and lots of monitoring and administration.
  • I don't want to write a single line of code that I cannot adequately unit/functional/integration test outside of the container.

What I've been tempted to do in this situation in the past is to just use Tomcat as my generic container, and model everything as a Servlet with no Servlet mapping, so I'm only getting the lifecycle calls that I need and use of web.xml initialization parameters as my configuration. But that just, well, it feels dirty. Plus, it would take longer than I think would be reasonable to launch one of these processor instances, and if I'm running thousands of discrete processor instances, I need control to be quick.

My guess is that there isn't something like this out there.

My guess is that I'm about to write one. I'd rather not, but I don't think that this type of thing exists. If it does, it might be the SpringSource Application Platform, but I can't actually tell what that does to be honest. I mean, it kinda just looks like OSGi++, but I can't tell what the ++ is. (Is it really just Equinox with some extra bits for central repositories and remote configuration? Is that all?)

My guess is that the actual eventual solution here is largely just going to be my leveraging Terracotta for all my distribution and failover requirements, but writing the whole actual lifecycle management work myself.

Oh, and it has to do Java. Telling me Erlang/OTP gives you every single thing here for absolutely free doesn't help me when I'm under a strict Erlang embargo at work.

Help me Lazyweb, Help me!