Friday, April 13, 2012

Maven pom tree



Here is a very nice trick I discovered almost by chance. Maven is an amazing tool to build software project. It scales...
Complex and big projects follows the SCM (Source Control Management) structure, usually a tree. Maven's best weapon is indeed to structure the project in a tree, where each node contains a pom that follows more or less the following rules:
  • module list is the list of children,
  • it shares definition with its subtree,
  • its parent pom is the parent in the tree.
  • its groupId is parent groupID plus parent artifactID
This way, at any place in the tree we have a good idea of what the effective pom is. The project groups together under a sub tree the type of artifact beeing constructed. Each level in the tree is refining the artifacts build instructions as the pom files provides more and more inherited definitions. For example:

documentation/
java/
-product1/
--web/
--wars/
--ears/
-samples/
tools/
-scripts/
--javascript/
--shell/
--ant/

etc.

Indeed, this is smart.

However sometimes in the tree there are duplicates. A java-like project can occur at several places and then it becomes impossible to share, except at the closest common parent.

Fortunately, there is a trick that uses a quite original feature: pom can refer to any arbitrary parent pom. The idea is then to keep separated a hierarchy of pom. These are very specialized and inherits all from a root pom.

poms/
-pom.xml
-pom-java.xml
-pom-war.xml

Now we keep the project tree, but when needed a pom can break the parent rule and refers to one of these poms instead of the tree parent. Different subtrees of same kind can now share the same parent pom.

Monday, October 19, 2009

The Turing Test in practice


While wandering on a merchant site, full of popups and flashing frames, my attention has been attracted by a new window popping up in the middle of the screen made of a title "New chat window", a text area which contains "operator>May I help you ?" and an input text field.

I was about to close this ad when a second line appeared "operator>Do you have a question ? I can help you."

This time I doubted. Fake or not fake ? I typed quickly "Are you a robot or a real person ?". The answer came a few seconds later "Don't worry, I am a real operator. May I help you ?".

The dialog went on and eventually the operator convinced me s/he was human. Not a very smart one, though. I was suggested to call the hotline for my request. But that's not the point.

The point is that for the first time I experienced the Turing Test. For a couple of minutes, I was engaging in a natural language conversation but I couldn't tell whether it was with an human or a computer, at least at the beginning.

For me, the Turing test is purely theoretical. We don't need to turn the test real until computers will be able to think. AFAIK it has never been experimented yet, except in science-fiction with its variant the Voight-Kampff test.

However, the conjunction of Internet, computerization of human communication (e.g. online chat), and creative ads pressure led to this situation. Not because computers are smart today, but rather because digital conversation brings human behavior down to the scope of a computer. But who knows, maybe that's the beginning of thinking computers anyway ?

Thursday, July 2, 2009

My favorite development process


After years of experience, I identified three kind of development processes that are commonly used for projects, according to their size, complexity, etc.




  1. Direct coding in a scripting language (like perl),
  2. writing code from requirements,
  3. writing code from specification documents.
Direct coding should be reserved for PoC and small projects, clearly doesn't scale, and is not suitable for teams. Requirements and specifications seem close, but actually it is a very different approach. The requirement gives no guidance on the implementation and comes usually from the marketing team. The specification is a very precise guide to the implementation but doesn't justify its content.

I've experienced the three modes, on volunteer basis or not, and eventually the one I prefer is... the fourth way!

What is the fourth way ? It's simply coding from the user manual. All the jigsaw that constitutes a project is supposed to take all its sense in the user manual. Start a project by writing the user manual is directly taking the customer point of view. The API is not only specified, it is explained. If a part cannot be explained clearly, it's likely there is a serious problem with it. And if it happens once the development is over, unfortunately it's too late...

Unlike the requirements, the user manual doesn't describe what the user needs, but what the user can do. The (functional) specification describes how the project behaves, so does the user manual which also tells how to use it.

When times come to write real code, the API and most of the comments are ready, which clearly accelerates the implementation step. The developer can focus on the algorithms and data structures. Unit tests can be prepared very early, too.

As usual it happens that unexpected implementation details has an impact on the user manual. Api must change to cope with a specific problem. The user manual is then updated. and coding can resume. Not a big deal, though, this kind of iteration belongs to agile development process.

In my experience writing the user manual first is very positive. And last but not least, as a side effect the documentation is always ready on time !



Saturday, April 4, 2009

The lost control flow


When I started learning programming languages, in Pascal, the emphasis was on the block structure. The advantages for the developer are obvious, like divide and conquer design, visible scope of variables, linear flow of control, etc.

It seems to me that things were getting too easy, and today the block structure is falling apart. The success of scripting languages, multi-tier architectures, and multiple layers makes it more difficult to follow the data along an application. Now the control flow is being attacked, too.

The multi-threading is obviously the first fighter which multiplies the control flows. But now there is an insidious one coming from event-driven applications. GUI toolkits widely use the listener pattern, because nothing is more natural than adding an event listener on a button so when the user clicks on it, a callback is invoked.

Unfortunately, the listener pattern is being overused, and events are more and more integrated to application logic, even where it is not really mandatory. The excuse behind is for example a cleaner separation between different modules, but triggering a callback disrupts the control flow and the external context is lost (which makes debug more complicated as the origin of event is less known and accessible). This wasn't an issue with the button, because the user click represents really a separate event that initiates a new processing.

Another cause for listener pattern abuse is the complexity of component life cycle. For instance, in Flex, the graphic component is build in several steps (creation, addition to container, measure, update to display with correct location). Flex 4 introduces extra steps for managing skins (skin parts are loaded, according if there are deferred, dynamic, or static). Eventually, the pre or post initialization of components are done through a mechanism like events, at one of the life cycle steps. Not only the control flow is lost, but the life cycle is far from being clear, giving the callback running some code although the state is not really the one the developer expected. For example, when a skin part is loaded, the size and location of the component is not set yet.

The control flow is indeed a cornerstone of programming. It's very sad the modern programming languages don't pay much attention to it and let it falling apart...

Tuesday, March 10, 2009

Ten years of my life...

I give away ten years of my life, to play guitar like David Gilmour! This song is a new version of the final solo of Fat Old Sun, an old but still one of the best of Pink Floyd. This version is available on the DVD remember that night, described by critics as a near-perfect gig.

I am able to work out Fat Old Sun on my acoustic guitar. Slow pace, 4 chords, tricky solo, very pleasant to play and hear, even from a beginner like me. But this rock'n roll version is totally different, and a master piece. That's why I need to invoke some divinities as a joker. So if the Evil is online, my blood is ready to sign the contract...

Monday, January 26, 2009

Teach it ...phenomenology


Imagine yourself in front of a talking Thermostellar bomb which is about to detonate, and start to verbally convince it to abort the countdown because, well, everything else failed...
This is one funny scene of first John Carpenter's movie Dark Star and one of the best (and extreme) dialog between a human and a machine. Of course, it's only a movie but the argument between Lt. Doolittle and Bomb #20 is not totally disconnected from our reality, I will come into it later.

Bomb #20 was stuck in its bay, about to detonate in the vicinity of the spaceship and its crew. It received the order to start the countdown, and is so enthusiast to fulfill its mission that a detail like still laying in the bomb bay is clearly not enough to abort its program. We can see that as a bug, but time is running out to debug. Lt. Doolittle's strategy consists in convincing the bomb it started the countdown on a false assumption.

That's why he used phenomenology, a philosophical branch that studies consciousness through the perception of external phenomenons. Bomb #20 cannot assert the countdown order was real or not. This question is critical, because the bomb can explode only once...

This perception/reality theme has been explored many times by Philip K. Dick, as well as strange human/machine interactions. In Isaac Asimov's books about robots, which happen to have also a highly developed A.I., Susan Calvin debugs robots using most of time philosophical approach. We can cite other movies or books where debugging process is far away from putting breakpoints in code.

Back in 2009, computers are not as smart and developers need low-level tools to debug software. But what happens in reality ? The developer knows it's useless to argue with computer. Therefore, the developer takes the computer point of view, reads the code and simulates (or plays) exactly what the program is supposed to do, while verifying if the state is normal or an error. In a way the developer argues with him/herself and navigates continuously between the two tasks, simulating computer and checking for fault. Errors can come from instructions, but also from bad data. The Bomb #20 argument takes place in the developer head.

But debug is not all. As programming languages provides more and more abstractions, the developer is confronted to higher level choices. For example, architects of Object Oriented applications often are facing semantic decisions. No doubt in the future the developer will need more reasoning skills that take inspiration from philosophical domains, starting from phenomenology.

Wednesday, November 5, 2008

Who cares about algorithms ?


The real question is: what is the value of an application. Is it the algorithm behind, or the user environment provided by the application. For most applications, the expected level in terms of performance and memory footprint reaches only "good enough". But what makes the difference is indeed the user perception of the application, that is the documentation, the online support, automatic upgrades, ergonomic, etc. In brief: how much the application takes care of the user.

In this perspective, the algorithm is secondary. It needs to be "good enough", since the user won't feel it directly. Years ago, it was different. Machine resources were most important to care about, and only a good algorithm could cope with them. The user interface was at his beginning and expectation in this domain stayed quite low. Today it's different, web technology asks for "instant customer care" and "attractive interface". For example, what makes a good web browser ? A fast and efficient html renderer or an optimized user interface ? IMO, I'm satisfied once the page can be be rendered in less than half a second and I start looking at how easy it is to manage and use bookmarks, plugins, and the overall ergonomic of the browser.

Note also that usually creating an algorithm is a job for one smart person, although it requires a full team to build up a good user environment, including tech writer, ergonomic people, QA, support, etc. Maybe it's difficult to find the person who is able to invent the best algorithm, but it's also very difficult to hire such an heterogeneous team and obtain a successful product at the end.

Friday, October 24, 2008

The lost Beautiful Programming Language

It was old time when a new programming language comes out, trying to be the best of all existing languages, designed by a Guru, aiming too fill a hole, etc. I think the last one was Java. Today when we see what a developer can do with an IDE like eclipse on Java, we really wonder what will the next language be and how can it be better.

However this is not taking account what mankind can do for its own loss. When we can't do better, we do worse.

The web pressure made plenty of "web" languages pop up in the landscape. One of them, Flex, looks very popular. But after playing with it, it becomes clear to me that Flex is backtracking.

Flex comes with Flex builder, a plugin on eclipse. However, the integration in eclipse is very minimal. No refactoring, no code assist, no automatic indentation, complex project settings. Only the editor, the debugger, a profiler, and the expensive wysiwyg designer that generates xml are available. I'm not sure it worths the 400M of my RAM the flex builder is eating. All in all, I miss Java support in Eclipse.

But IDE support is not all. Flex syntax is confusing, especially when it mixes xml and ActionScript. Even the semantic is misleading. Let's take an example:


var t:Object = { foo:bar } ;


The right hand part creates an object, actually an "association list". It contains a property of type key/value, where key is "foo" and value is "bar". "foo" is a String literal, but "bar" is supposed to be an object. The value is stored in the association list, but is eventually converted to a string. To use it:


t.foo ;
// or
t["foo"] ;


Note the very unusual use of "." and "[ ]". Anyway, this construct is very handy. Basically, an association list is an hashtable.


I let you imagine how much time it needs to figure out this simple feature from the documentation, and how error-prone can be a language with plenty of this kind of counter-intuitive syntax. I know, this is common with scripting language. But can we still speak about scripting languages when it has packages and an object-oriented layer?

So yes, Flex is the royal language for RIA, but no, it's not as slick as Java.

Tuesday, October 14, 2008

I fixed a bug in 37 minutes


Obviously, this is not an outstanding performance. The statement looks very common, however it assumes a lot of things:
  • I was able to measure the time it really takes. Actually I use mylin extension of Eclipse to help me drive my work by task. This time, I don't know why, I paid attention to the elapsed time to completion, a built-in feature of mylin I never used before.
  • I was not disturbed or distracted during 37mn in a row. It seems to me it has been years I couldn't focus so much time on a single task.
  • The bug itself wasn't a big deal, but I think I am the only one on earth to fix it in such a short delay. It's not because I'm super-developer, but because I know the best this part of code, and my developing environment is ready to tackle this kind of bug. Even a genius challenger would need a couple of dozen of minutes only to setup the right environment. Actually the difficulty of this bug was to find out the right place to fix it, and I used the debugger to find this place quickly and accurately, instead of navigating through the sources. The fix itself was simple.
  • I searched in google code source an implementation of a method that was part of the solution. I could implement it myself, but it saved me a bit of time.

It's only unfortunate that the time I saved fixing this bug has been wasted writing this post!

Friday, September 19, 2008

Sofware equations

I've told already about:


Software = Algorithms + User Interface + Bugs


Here is another one:


Programming Language = syntax + semantic + doc + libraries + IDE


The two first, syntax and semantic, are mandatory to define the language. The other ones are simply making it respectively understandable, useful, and usable.

The documentation contains the reference manual, the user manual, tutorials, and sometimes other kind of document that helps the transition from another language to this one. Ususally, a language is not coming up from scratch. It reuses known concepts (e.g. object oriented, multithreads) and can be very close to another existing language. For instance, Java syntax is similar to C, so dedicated books explaining Java for C programmers are very popular.

Libraries are making today the differences between languages. The object oriented paradigm spreads the concept of library driven by an API. And the web technology is a great producer of layers and libraries. As more and more libraries are integrated, the language becomes more and more useful.

IDE is now inevitable. Once a developer has tested a serious IDE like eclipse or visual C++, it's very difficult to go back to antic era where code is written in a text editor, compiled, debugged with an external debug tool (when it exists), and deployed through a shell script. The IDE provides the editor, compiler, and debugger, but also it suggests code patterns, completion, syntax highlighting, online documentation, and provides powerful navigation, packaging, etc.

Wednesday, July 2, 2008

flex / silverlight comparison

I've just attended a presentation comparing flex and silverlight. In short, both languages are equivalent in terms of functionalities and libraries, which is not really a surprise... However, the noticeable differences are:

- IDE: Flex builder exploits 10% of eclipse capabilities and is far behind visual studio (but is far less expensive). There is not even automatic indentation!
- performance: current silverlight version (2 beta) is 3 times slower than flex 3! Remember flex is itself 3 times slower than Java.
- deployment: flex plugin is deployed at 97%, silverlight is probably closer to 3%.

Of course these differences holds at this very time. Both platforms try to catch up with the other to become the leader, but no doubt flex is ahead for the moment.

Friday, June 13, 2008

Software = Algorithms + User Interface + Bugs

This point of view holds for desktop applications, but also for any embedded software or more generally, everywhere there is a programmable chip.

The algorithm part is the most obvious for developers. It represents the processing of data and more recently the integration of libraries at different levels to provide compatibility between the various operating systems, languages, and protocols (read for instance: the web is a mess). Algorithms are the engine of the software. When developing resources are low the temptation is high to focus only on this part, because it produces the main output.

User interface is the steering wheel and the dashboard of the software. It becomes more an more an issue especially for portable devices, like a mobile phone or PDA where it really makes the difference between a useful gadget and a piece of crap. A bad user interface cancels all efforts provided by the software and hardware behind. UI is not trivial. The communications between machine and human goes mainly through a screen in one direction, keyboard, buttons, and mouse in the other direction. This is not always intuitive, and that's why UI is critical as machines become more complex.

Unlike Algorithm which is an hard science backed up by the power of mathematics, UI is an cognitive science. There is no formula for UI, only general guidelines for developers and eventually statistical studies bring the most reliable metrics.

Bugs are the dark side of the software. Nobody wants them, they are here. They are just part of the process. In our car metaphor, the bugs are all what don't work as expected from the annoying inside light that don't work when the door is open, to what can cause or aggravate an accident.

UI is a cognitive science, bugs are no science at all... They are unpredictable, and apart research projects around formal provers that are applicable in specific conditions, only intensive test suites can find out the bugs in the general case. But it remains the environment around bug fixing. I mean if we accept software have bugs (and we should do), then let's spare resources and prepare tools to handle them efficiently once they are found.

To take care of bugs, a dedicated team named QA is needed. If QA has a price, the lack of QA has a much higher cost. For most secure applications like embedded software in planes, QA can represent up to 70% of the product price...

The common point between Algorithms, UI, and Bugs is that they are much more easier to handle when they are taken account early in a product life cycle. It's not once the product is out that you start taking care of UI or tests ...or the algorithms!

Tuesday, June 3, 2008

OS scheduler and grilled beef skewers

BBQ
Did you try to cook beef skewers on a BBQ ? A BBQ is not a balanced heat source, and skewers are difficult to turn around. The result is that some pieces are burned while others are still raw.

For sure OS makers are aware of this problem. The proof ? Try to run an infinite loop in a shell, like in bash: while true; do set x 1; done, and look at the task manager to see how busy are your multi-core machine.

The naive answer suggests one core is 100% busy while others are free. However the reality is different. On XP, the OS scheduler happily migrates the infinite loop process from one core to another making all cores partially busy with this infinite process.

Advantages: the only one I see is that load balancing avoids extra heat on one part of the CPU, exactly as if the skewer where regularly turned and moved around all over the grid to be better cooked.

Drawback: the process migrates, which means in addition to context switch overhead the data are copied from one L2 cache to another. Overall time is longer than on a mono-core machine.

Workaround: You can pinpoint a thread on one core and prevent it to migrate elsewhere. This is called "single core affinity".



Tuesday, May 13, 2008

Multi-threading: the mine field is ahead

You think you master multi-thread programming ? CPU makers are asked to improve CPU efficiency, without paying
attention to the software behind who needs to cope with the new features. To put it clear, CPU makers are working on your next traps.
  1. memory access swapping
  2. instruction reordering
  3. asymmetric cores and unfair memory access

A read instruction is a lot faster than a write, and the memory bus cannot perform in both directions at the same time. Memory access, even a read, is still slower than CPU basic instructions. To keep the CPU busy, it may be needed to swap some memory access. For instance if the CPU has enough data in the cache to run without other memory access, it's a good time to perform an expensive write, even if read access are scheduled first. Also, some instructions can be reordered to optimize the processing pipeline and take advantage of the co-processors.

These arrangements inside the CPU are managed by a code analyzer which detects commutativity of sequence of instructions, and independence between variables. This analyzer works in a single thread context, so it's up to the developer to identify multi-thread issues and to forbid the CPU to arrange some parts of the code by synchronizing code and protecting data access.

Examples are numerous, here is a simple one, in a symbolic language

boolean ready = false;
int result = 0;

thread1:
while (!ready) {
Thread.yield();
}
print(result);

thread2:
result = 11;
ready = true;



The human understands quickly the purpose of this program. Thread2 is producing a result, toggling a boolean when completed. The other thread waits for the result to be ready before consuming it. But without any protection, this program can print 11, 0, of even nothing because thread1 loops forever!

That's because if you imagine yourself a single core executing thread2, you don't bother writing back the values result and ready because you don't need them afterward. Only a context switch forces you to flush the cache and then to update the memory location of these variables.

The core running thread1 is even worse. ready and result are not bound by an expression. They are independent so they can be fetched in any order. In particular, if result is read before ready, it prints 0. Also ready, once loaded on the cache may never been refreshed from the main memory, which leads to the infinite loop.

Sometimes it runs just like human expected, and the program prints 11. But that's a lucky execution, actually...

Last mine in our field, for the moment, is the asymetric architecture: a core has some privileges, runs faster or has more cache, or gets higher priority to the memory bus. It's a nice idea. An important thread can run faster than the other. But once again the software is far behind. OS kernel are usually SMP (symetrical multi-processor), and dispatch processes in a fair way across the processors, which is exactly what we don't want. Same for the higher level.

Assuming we have the possibility to pinpoint an important thread in the best core, the developer still have to think again the multi-thread logic to optimize the core assignments. Still good development time to go...

Friday, April 18, 2008

RIA: is Java out ?


I've just read Hybridizing Java from Bruce Eckel, the author of Thinking in Java. Bruce was an early adopter of Java and now he opens fire on the language. How could he change his mind so radically ?

Thinking in Java was my primary book to learn Java, 10 years ago.
I still recommend it to my students or colleagues as it is ideal for developers with already a good knowledge of computer languages. And it's online, and free (all but the last edition, though).

Hybridizing Java is puzzling me. Some statements are obviously wrong or inflated (e.g. more and more sites are not compatible with Firefox, Flex is the only plugin that solves nicely the UI question on the browser side and without installation hiccups). Being a Flex evangelist (and Adobe consultant) shouldn't allow to be so rude with what made him so famous before.

Besides that, some thinkings are right. If a language didn't address correctly some issues in ten years, it's normal users are getting less confident and look elsewhere. This is what happens with RIA (Rich Internet Application). Java applets should dominate the domain, but hasn't. Flex takes it over. Even if Flex is not as powerful as Java today, it may become in the future. At least it allows to create attractive RIA. Flex is good enough, and the quantity of sites that use flex is definitely a proof.

"make simple things easy, and difficult things possible" is a fundamental Java design guideline. Flex simply did it better than Java for RIA.

Friday, March 28, 2008

When guitars meet computers


I am always surprised by the ratio of amateur artists among engineers. For example, we are 5 guitarists among my 10 closest colleagues in the office.

I guess it's likely the need to balance rigid computer logic with forgiving art. Music is a frequent choice, guitar seems the winner over piano (because of yet another keyboard?).

But the interesting point is that one kind of guitar, the electric guitar, wakes up the computer geek when he plugs the guitar into the microphone outlet of a basic sound card: it works.

The geek has just entered into a new kingdom, the DSP (Digital Signal Processing) land. There are tons of software, mainly VST plugins, that simulate digital effects and turn a common PC into the equivalent of hundred of kilos of racks of hardware and kilometers of cable. As described in this page, in French but with a lot of images, the software is rich of attractive GUIs with buttons, sliders, and visualization gadgets. All you need is a computer, a basic sound card, decent speakers, and the software. With very few investments the result is impressive, because the sound is really great.

The geek is now ready for his first quest: the Perfect Sound.

Once he's satisfied with the sound, let's aim to the second quest: the content. Here is the second advantage of the computer: internet is a huge repository of songs, guitar tablatures, guitar lessons, and even video of guitar players.

Then it's not very fun to play alone, and here again the computer helps: it provides an orchestra, playing mp3 or midi songs along the guitar.

And here I take my revenge over the computer. It rejected my program because I forgot a semi-colon, I impose it all my rehearsal with always the same faults at the same place. And when it's finished "play it again, Sam". For once, it follows all what I want it to do.

More precisely, I'm specialized in Pink Floyd's solos, like Is there anybody there, Time, or Fat old sun. I'm very impressed by David Gilmour plays. The solos are usually slow-paced, with few notes, but each one sounds great and contributes to a beautiful harmony... He uses a lot of bends, which consists in pushing the cords on the fret to apply extra tension and then augment the note pitch gradually. Plus very small tempo shift, it gives a solo full of tensions that drag brain attention, followed by a relief as the play catches up back to normal pitch and tempo.

Gilmour's solos looks simple on the paper, but believe me there are very hard to work out with the same touch. Anyway. I'm having a lot of fun with the electric guitar plugged into my pc, the sound presets, the midi orchestra, ...

Thursday, March 20, 2008

The most expensive bug


1996 june 4th, the first european rocket Ariane V exploded 40 seconds after launch. The payload alone cost about US$370 million. The cause ? A bad cast in the software initiated a chain of dramatic errors and led to Ariane destruction.

The full report worths the reading, here is a summary.

The attitude of the rocket is given by an Inertial Reference System (IRS), which is a combination of gyro lasers and accelerometers. This critical piece of hardware sends a stream of data about position, height, speed and acceleration to the main computer, which controls the exhaust pipes and drives the rocket along its expected trajectory.

Ten years ago, on Ariane III, a software function performed pre-flight checks to test alignment of the IRS. This function was no longer used on Ariane IV, but still ran during take off. You know it's easier to leave harmless code than to remove it. This function used 8 variables, 3 of them were not correctly protected although it was not an issue, because the rocket trajectory remains in range of these 3 variables.

No surprise, this function was still working on Ariane V. Unfortunately, Ariane V trajectory was a bit different and now one of the variables, the horizontal velocity, casted from 64bits float to 16bits integer, went out of range and raised an uncaught exception.

So far, no big deal. A check function raised an exception. Let's forget the check function and resume the mission.

However, the assumption on Ariane design was that software is always right and hardware may fail. The software reported an error, interpreted as the SRI was out of order. Then the SRI was shut down.

That's probably the biggest mistake. A failing unit test is embarrassing enough, but doesn't always mean the software is out of business. In this case the SRI still delivered reliable information. Unplugged, it couldn't any more.

The backup SRI started providing replacement data, and was shut down 0.05 second later, because of the same bug. Once again the assumption "hardware may fail, software not" made the backup SRI totally useless in this case.

Without sensible guidance, the rocket was doomed. But to accelerate the disaster, the SRI modules started to send stack traces instead of normal data to the main computer. The computer interpreted the data just as if the rocket was upside down and went into an emergency half turn. It started to tear apart under the physical constraints and initiated self-destruction process.

The story is sad enough like that, no need to add that a suitable test or full simulation before the flight would have found the bug.

By chance, I knew one of the member of the investigation team. He told me something not in the final report: what greatly contributed to kill Ariane V is the absence of experimented computer scientists at top management. The sofware components were simply divided and individually conducted. A competent software supervisor with suitable power could have found one of the errors, and prevented the cast exception to eventually stop the delivery of correct IRS data.

But Ariane was a physicist toy, they didn't share with a software department...

PS: The lesson was positively received. Today Ariane V is very successful and crashed only once more in 37 flights.

Friday, March 14, 2008

Marcel-Paul Schützenberger and complexity

It may happen in your lifetime that you have the luck to meet extraordinary people. MPS is one of them. I attended his lectures when I was a student at the university (Paris 7). This guy is not really famous, he didn't look for fame. We were always less than 10 in the audience. So who is he ?

MPS is a physician, a mathematician, and a computer scientist. Obviously he is excellent in all three domains. His various knowledge gives him a very realistic vision of computer science. Basically, computers were for him a tool to develop mathematics, and have a great future in biology. The last part sounds obvious today, but he said so more than 20 years ago.

But what makes him so attractive during the lectures is that he has stories to tell. As a founder of modern computer science, he met all the other founders all over the world (ok, mainly in USA), invented a very important theorem in language theory, and so he has a lot of nice anecdotes to tell about this pioneer period. I barely remember a couple of them I keep for another blog entry.

For the moment I want to focus on a sentence he said, which is carved in my brain forever:
"There are 2 kinds of program. The short ones and the long ones."
This can be understood different ways. I think the basic idea is that we should keep our distance from computer power and program complexity, and whatever we are trying to develop, after all it's only a computer program, so let's just break it down in sequence of instructions.

For example in the 80's, the hype was on Artificial Intelligence. For a majority of people AI was the most complex applications we can even think about, so much complex that nobody could complete them, btw. Eventually AI died because it didn't fulfill its promises. Some blamed computer performance, although computer power doubles every 18 months it will never catch up with AI complexity which scales up to infinity. Other blamed the poor expressivity of computer languages, too low level. That's better, but no. The real reason that killed AI is the lack of theory. Without theory, no suitable language and no proof of algorithm termination. Then your computer program can be as long as you want, it won't implement correctly the specifications (or you won't be able to prove it). Conclusion is that the program doesn't make it all by itself, it's only a tool and it won't cope with a hole in the theory.

Note that MPS didn't want to under-evaluate complexity in programming. For the developer, the complexity is relative to the constraints s/he has to deal with, about the programming language, tools, software architecture and design, and available resources. But the program is only the implementation of an algorithm, and a correct and efficient algorithm relies on a strong theory.

Friday, March 7, 2008

The Next Programming Language



Everybody knows Moore's law: "Computer performance doubles every 18 months". But programming languages have also their growing law: "A very new language appears every 10 years". "very new" is indeed relative, and should be understood as "language with new features and successful". Let's review the last ones:

  • 1972 : C was the first high level language bound to an operating system (unix). For the developer, it means a very fine grain control on the host machine and the ability to use a high level language to program at low level.
  • 1983 : C++: C with an object oriented layer. Note that one of the most claimed feature was the portability, which happened to be a disaster (C++ libraries were less compatible than C ones). C++ went to far. Macro, ability to override the operators, multiple inheritance made eventually the applications a nightmare to maintain and to integrate.
  • 1996 : Java: Object oriented, native multi-thread, GC, beans, exceptions handling, no macro, no multiple inheritance, clean packaging, portability, applets as the very first browser plugin, etc. A lot of advantages.
  • 200* : Web scripting languages. javascript, php, flex, XUL, etc. We can't say that one is the leader, but for sure they all contributed to empower the web and to make the web experience as sophisticated as a full fledged application. Some people says it's rather .NET/C#. Sorry, I disagree. There is no revolution with C#. It stays in between.
  • 2010 : What is waiting for us ? My guess is a new language that will support multi-thread as simply as Java managed the memory for the developer.
Multi-threading is really calling for a new language. The multicore architecture need a fine grain control over thread dispatching across the cores. When 2 threads are expected to communicate a lot, they should be running on two close cores so they can share the same L2 cache.

Most of multi-threaded languages, like Java, offers synchronization by locks which is probably simple to implement on the OS, but surely the most difficult for the developer. I really believe there are other viable solutions, like continuous transaction already working for database, which reliefs the developer synchronization logic and all race condition/block/starvation/CPU contention bugs. These bugs are a pain to track down and fix. We have almost no tool to help, and no background theory. Sigh.

Java is my favorite language, but I must admit it's out regarding multi-thread. The new JDK1.5 package java.util.concurrent helps a lot, but basically it doesn't get rid of the complexity.

Beside that, the next language will look very similar to Java, with smart packaging, Object oriented layer, GC, etc., and usual syntax for control statements.

Thursday, February 21, 2008

Multithread: the new era


Multi-threading is not new. This is, for instance, natively supported in Java since the beginning in 1996.

What is new started with this very popular article that appeared in Dr. Dobb's Journal, 30(3), March 2005. Basically the heat produced by the CPU doubles when the clock is 20% faster. Usual fans and radiators are saturated with CPU around 4GHz. It means the clock race is over. But processor manufacturers are smart. If they slow down the CPU by 20%, it produces half the heat, so let's put two of them in the same chip! The dual core is born. Same thermal envelope, 2 cores, 20% slower than mono core version. Moore's law still hold if we consider each core as a contributor of the total machine speed.

By the end of 2007, almost all new machines are equipped with multi-core processors.

From the software point of view, it has 2 major impacts:
  • Customers are now running your applications in true multi-threaded environment.
  • The multi-core hardware speeds up performance only if the application is multi-threaded.
We are still correcting bugs coming from first point. Even Java made mistakes, for example the famous single thread rule of Swing happens to fail on multi-core machines.

We are only starting to tackle the second point. And we have very few tools and theory background to help us so far.

After the Recursivity and the Object Oriented Languages, now comes the Multi-Threading.