Thursday, February 3, 2011

Mondrian SPI SegmentCache

Fellow Mondrian developers and users,

One month has already passed since the new year festivities, and while most of you have been trying to renew your gym membership or hold on to your new year resolutions the best you could, so did the Mondrian team. Our resolutions, although not requiring personal sacrifices, are none the less starting to bear fruit.

For you see, our resolution for the year was to provide Mondrian developers and integrators means to achieve better understanding, scalability and control. We have many ideas on how to reach those goals. Some of them are still in their infancy, yet some of them have already been committed to the source. Last month, we worked on the first phase. We added means for system architects to externalize and share a  pluggable segment cache. What does this mean exactly? Let's take a step back in order to better understand.

Internally, Mondrian splits the tuples in segments. A typical segment could be described as a measure crossjoined by a series of predicates. As an example, a textual representation of a segment contents could be:
Measure = [ Sales ]
Predicates = {
          [ Products = * ],
    [ State = California ],
    [ Gender = Male ] }
Data = [ 1346.34, 234.00, ... ]
In the case above, the segment would represent the Sales data of all males in California, for all products. It is a lot more effective to deal with those data structures. If Mondrian was to internally represent each data cell individually, the unique identifier of that cell would be of a greater size than the data itself, thus creating a whole lot more problems in terms of data efficiency. This is therefore why Mondrian deals with groups of cells, which it loads in batches, rather than individually. There is a lot of voodoo magic and heuristics in the background trying to figure out how best to group those segments and how to reduce the number of segments to load, ultimately reducing the number of SQL queries to be executed. Mondrian will group all segments with the same predicates but with a different measure into a segment group. Mondrian will also tend to remove as many predicates as it possibly can in order to optimize the data payload. Lets say that a segment covers all products except a single one, Mondrian will still include the product in the segment but filter it out when a specific query requires it.

Once those segments are populated, Mondrian keeps those in a collection of weak references in local memory. All required segment references are pinned down during the resolving of a particular query, but as soon as the query is done executing, the references are returned to their weak state, thus ready to be garbage collected if needed. This simple mechanism allows Mondrian to answer just about any query, as long as the memory allocated is big enough to answer that particular query. This works really well in fact, since in most small deployments, the maximum amount of memory is never reached. And if it ever gets filled, old segments will be evicted to make some room for the new ones.

Now, there are obvious gotchas. First off, what if it takes a long time for a segment to be populated by the RDBMS. This means that if a particular segment ever gets picked up by the garbage collector, the MDX query sent to Mondrian *might* take longer to execute, whether it was in the segment cache or not. This is not acceptable, simply because this makes all performance predictions impossible.

This is where the SegmentCache SPI comes in. It is essentially a pluggable cache for segments. The algorithm behind the segment loader becomes this:
  • Lookup segments in local cache and pin those required.
  • Optimize / group segments
  • Lookup segments from the SPI cache
  • Load the segments found from the SPI  cache
  • Populate the remaining unloaded segments from the RDBMS
  • Put the segments which come from the RDBMS into the SPI cache
  • Pin all loaded segments
  • Resolve the query
  • Unpin all segments in the local cache
But wait! There is more! The SegmentCache SPI is trivial to implement.
Future<Boolean> contains(SegmentHeader header);
Future<SegmentBody> get(SegmentHeader header);
Future<List<SegmentHeader>> getSegmentHeaders();

Future<Boolean> put(
  SegmentHeader header,
  SegmentBody body);
void tearDown();
Figure 1. Mondrian Segment Loader Architecture


There are two assumptions that are made towards the implementation. The first obvious one is that the cache must assume that many Mondrian instances might access the cache concurrently, form different threads. We therefore recommend using the Actor Pattern or anything similar in order to enforce thread safety. The second is that SegmentCache implementations will be instantiated very often. We therefore recommend using a facade object which relays calls to the actual segment cache code. Update: This was redesigned so that a singleton is created and used throughout Mondrian's internals.

As for the storage of the SegmentHeader and SegmentBody objects, we tried to make it as simple and flexible as possible. Both objects are fully serializable and are immutable. They are also specially crafted to use dense arrays of primitive data types. We also tried to make extensive use of Java native functions when copying the data to / from the cache within Mondrian internals.

The bottom line is that from now on the Mondrian community will be free to implement segment caches to fit their needs. We will be rolling out a few default implementations and examples, obviously. One neat implementation could be one which pages the segments to a super fast array of SSD drives. Another one could be to store the segments in Terracota or ehCache or Infinispan, or just about any scalable caching system there is out there. So if any of you out there are interested in implementing this SPI for your business and would like to either share your experiences or contribute those implementations, don't hesitate to contact us. Or me directly.

There is more goodness to come, but that's it for now. Stay tuned!

Monday, November 1, 2010

olap4j 1.0 - The long road to LTS

Last summer, we at olap4j announced that we would release olap4j 1.0 on October 31st 2010. Now is November 1st and olap4j is still not out of the door. Here's why.

Our reference implementation, the Mondrian project, suffered from delays in the release dates and had to be pushed back. All this kept us very busy and scrambling to make the best of it for a while. But fear not! This very week, the Mondrian team is running QA tests on a 3.2.1-GA build and as soon as it gets the green light, we will be able to put the finishing touches to olap4j 1.0. Why are these projects tied so closely? one might ask. Legitimate question.

You see, every API is a better API if there is at least one reference implementation existing in the wild. For many reasons actually. It allows us to architect, develop and test in a real environment, with real constraints and real data. Several APIs have failed in the past because in the end they were overly complicated for the end user. Some were so overly complicated that in the end, they failed to deliver what an API is supposed to be; a simple interface to a given system. Having a reference implementation mitigates the risk by allowing us to release by little increments and release often, but most importantly, release testable code that works. Another huge advantage is that for every part of the API, there is at least one fully functioning implementation out there, freely available as open sourced software, which implementers can refer to. This is a huge advantage in terms of both project sustainability and project adoption.

So stay tuned, because olap4j 1.0, despite some delays, is right around the corner!

Tuesday, September 21, 2010

olap4j tutorial

Behold! I finally had some time to put the finishing touches to my tutorial for olap4j. It took a year but here it is. Sorry for the delay. I know, I know. I promised it a year ago, but a lot of stuff has been going on since then. Two new jobs, two moving to a different city (and once more next month...). So without further delay, enjoy!

http://olap4j-demo.googlecode.com/svn/trunk/doc/Olap4j_Introduction_An_end_user_perspective.pdf

Wednesday, July 21, 2010

From the olap4j team

I posted this message today on the olap4j mailing list. In the interest of reaching a broader audience, I will copy it here as well.

Cheers!

Dear olap4j community members,

As we previously discussed on this mailing list, we are planning to make the final push towards a 1.0 specification. In order to perform those much needed changes and still maintain compatibility as much as possible, the olap4j team proposes the following transition plan.

  • 4th week of July - Release of olap4j 0.9.8
    A first initial release, coded 0.9.8, will be performed during the days to come. This release is mostly a wrap-up of the unofficial releases we have put in the Maven repository. Notable changes include compatibility with SAP BW, contextual drill-through for the Query Model, along with various other compatibility fixes.

  • Month of August
    During the month of August, the olap4j community will have a chance to comment the proposed changes to the 1.0 final draft. We will provide an updated functional specification document as well as a complete list of the changes that will be required. Should you judge that some items are still missing, or that some should be modified or removed altogether, you are encouraged to let us know. The mailing list is the best place to hold those discussions, or you can also use our forums.

  • September 1st - Release of 0.9.9
    September 1st is the date that marks the end of our discussions. After that, all the changes that we agreed upon will be implemented in the API, as well as the Mondrian and XML/A implementations of the driver. The 0.9.9 release will include those changes, but will still maintain retro-compatibility. Some API calls will be marked for deprecation, new ones will be present as well. The 0.9.9 release will be the last available before 1.0. Everything that is marked for deprecation will be removed as of 1.0, so users will have a chance to convert their code base progressively.

  • October 31st - Release of 1.0
    We are planning to release olap4j 1.0 on October 31st. All methods that have been marked as deprecated, whether by the 0.9.9 release or any other previous 0.X release will be removed.


A proposed updated functional specification document, as well as a detailed list of API changes will be sent in the following days, right after the 0.9.8 release. We strongly encourage the users of olap4j to express any concerns or ideas that might arise.

Sincerely yours

Luc Boudreau, for the olap4j development team

Tuesday, June 29, 2010

Pentaho’s Road To Profitability - Take 2

Tom Barber, a collaborator of mine, one of the most active community member of the Pentaho project, posted a very interesting blog entry this week. He is looking back at the path of Pentaho Corporation, a commercial open source business intelligence company, and their latest strategies for growth. He notes that in the past months, Pentaho has put a lot of effort in marketing initiatives and hired many big wigs in their marketing staff. In his opinion, this is somewhat against the ideals of commercial OSS development.

The commercial open source business model is still very young and has not encountered many great challenges up to now. Nor has it sailed in very troubled waters. Yet some signs are already warning people to be very careful with the years to come. Sun Microsystems (now Oracle) was surviving thanks to donations. Compiere recently made the news for all the bad reasons. Red Hat is doing pretty well though. In a nutshell, anything is still possible; good or bad.

Tom’s wish was that Pentaho would rather focus on paying skilled engineers rather than sales people. Now, as much as somewhat agree with the general idea, one must keep in mind that there is no correlation between the number of talented people getting paid to be on a project and it’s success. Some notoriously successful projects depend almost entirely on it’s community base. Mozilla Corporation to name only one. Others employ thousands of employees, yet achieve mediocre results. One could also argue that an effective marketing campaign will in fact boost people’s awareness, thus getting more talented people to join the community base. There is no tested and fail proof recipe so far. As I said earlier, everything goes.

I worked for the past months for a commercial open source company, and the same questions and uncertainties were part of every day discussions. How can a software company making no revenues on licensing be profitable? SQL Power Group, my employer, sponsors the only open source data modeling tool that works cross-platforms and offers, for free, the majority of the functionality included in widely known proprietary tools like ErWin. There are on a weekly basis between five hundred and a thousand downloads of SQL Power Architect. This is a lot for an OSS project in such a narrow niche. Yet a very small proportion of those are actually paying for support/consultancy, or making donations.

There are a lot of factors in play. First off, people willingly using OSS have overall better technical skills than others. To be fair, not every OSS is of acceptable quality, but you get to try as many as you want. Proprietary software is not better in any way, but you won’t get to try many of them, short of having very good contacts and/or deep pockets. This is one reason why reaching the tipping point of adoption is very hard for an OSS company. The percentage of OSS users ready to pay instead of figuring things by themselves is low. Way low. Waaayyy loooww. If you paid for software, you want support (and vice versa). You feel obliged to use the software because you already invested much in it. (Yes, money is time, time is money.) If you didn’t pay for the software and you hit a snag, you are far more likely to stop using it altogether. By abandoning it, you feel like you are exercising damage control, since you save time (thus money). You didn’t pay so why invest time? I’m fully aware that once you do the math, you realize that it’s the exact opposite. People usually don’t.

What does all that have to do with Tom? Well first thing first. Starting next week, I’ll be an employee of Pentaho Corporation. Yup. Am I a big wig? Nope. Am I a marketing guy? Well, I’m good looking, but not good enough with oxymoron. Am I a marketer? I’m having trouble making this blog post interesting, so nope, not a marketer either. I’m just some skilled dude who got his hands dirty for years and now decided that he would devote some serious time to make things better.

Look at it this way. In Tom’s vision of Pentaho, my ass is on the line. I deeply respect his insights on the market and software development in general. Yet here I am, proud as a peacock on a Sunday brunch, about to be part of this fantastic project. Am I worried? Hell no. I’m very enthusiastic about Pentaho having enough budget for marketing initiatives. This is exactly what OSS companies need; market awareness. People need to know you exist and that you are just as good as all those Fortune 100 companies. Even better than these. It has nothing to do with being a sell out. Marketing is not bad for your grassroots values, unless you make it so. Then again, you have only yourself to blame.

OSS companies need to reach the critical mass of contributors and adoption needed to survive. There is simply no other way. OSS users are picky, grumpy, bitchy and unforgiving. I know that for a fact. You want to thrive in that market share? You need two things. Proper marketing and skilled engineers. I’ll be doing my part in the later, and I’m very confident that the good people at Pentaho have picked skilled people to cover the former.

We are competing against giants, so let’s give them a ride for their money.