Friday, December 27, 2024

Getting Started with AI: NotebookLM

 If you followed along with a previous post on a holiday challenge for learning AI you may now be wondering where too next? Great question, shows you have learned about prompt engineering and are now thinking there has to be more. There is more, a lot more. A good set of skills and understanding of prompt engineering would serve you very well, and you could stop there for a while. Particularly, if you iterate your prompts and increase your literacy in creating prompts. And remember AI can help you improve your prompting.

For many people, I have found that once the intermediate understanding of prompting is achieved the question doesn't seem to go to how do I prompt better. The question seems to go to, can this be automated or can my LLM be more subject specific or can the LLM be restrained personally to my own knowledge. I want the AI to be more specific or give more weight to a narrower or more personal domain of knowledge.

Retrieval-Augmented Generation (RAG)

Retrieval-Augmented Generation (RAG) is like giving an AI system a personalized library to reference while it's talking to you. Instead of only relying on what it learned during training, RAG lets AI search through specific documents or data to find relevant information before generating a response. Think of it like a student who first checks their textbook and notes before answering a question, rather than just going off memory. This helps the AI give more accurate and up-to-date answers based on reliable sources. 

There are a few online options to provide you a personal RAG platform. Currently, my two favorites are perplexity.ai and NotebookLM. Both these platforms allow you to upload or reference other resources (text, video, and others) to augment (and focus) your use of AI. Really very amazing at supporting you in creating subject specific AI mentors. I strongly suggest you begin to play with NotebookLM.

Consider using NotebookLM

  1. Set up an account (or use it with your existing google account). https://notebooklm.google.com/
  2. Watch a NotebookLM introductory overview video: https://youtu.be/UG0DP6nVnrc?si=2bGoT7ZMI-VKsU6_
  3. Think about business and personal use cases: https://youtu.be/U3SgtCWsjXg?si=eR_ESarUJTHenPki
  4. Consider the history of NotebookLM development at google: https://youtu.be/sOyFpSW1Vls?si=F9gVrxXrc2vihRnf
If you remain curious about where RAG and automated agents fit into all the near future of AI I published a post last week discussing these two innovations with AI. Twenty twenty-five will be an interesting year.

Wednesday, December 18, 2024

RAG and Agents: How AI is Learning to Think and Act

Collaborative RAG and Agents.
In the rapidly evolving landscape of artificial intelligence, two technologies are fundamentally changing how AI systems interact with the world: Retrieval-Augmented Generation (RAG) and AI Agents. While both enhance AI capabilities, they serve distinctly different yet complementary purposes in advancing machine intelligence.

RAG: The Power of Grounded Knowledge

Imagine trying to navigate a foreign city using only your general knowledge of how cities work. You might make educated guesses about where to find the downtown area or how the transit system operates, but you'd likely make many mistakes. This is similar to how traditional Large Language Models (LLMs) operate – they rely on their training data to make informed but potentially inaccurate assumptions.

RAG transforms this paradigm by giving LLMs access to specific, relevant information in real-time. Instead of relying solely on their training data, RAG-enabled systems can pull precise information from your organization's documents, databases, and knowledge bases. This means when you ask a question about your company's Q4 2023 results, the AI isn't generating a plausible-sounding response – it's retrieving and synthesizing actual data from your financial reports.

The impact of RAG on accuracy and reliability cannot be overstated. In healthcare, for instance, RAG-enabled systems can access the latest medical research rather than relying on potentially outdated training data. In legal applications, they can reference specific case law and regulations rather than generating generic legal-sounding language.

Agents: From Knowledge to Action

While RAG revolutionizes how AI systems access information, Agents take things a step further by adding autonomous action to the mix. An AI Agent is more like a capable assistant than a simple question-answering system. It can:

  1. Plan and execute multi-step tasks
  2. Interact with external tools and systems
  3. Maintain context across conversations
  4. Learn from past interactions
  5. Make decisions based on evolving situations

Consider a customer service scenario. A RAG-enabled system might accurately answer questions about your return policy by referencing your documentation. An Agent, however, could actually process the return, check inventory for replacements, schedule a pickup, and update your CRM – all while maintaining a natural conversation with the customer.

The Synergistic Future

The real magic happens when RAG and Agents work together. Imagine an AI system that can not only access your entire corporate knowledge base but also take action based on that information. It could:

  • Monitor market trends and automatically adjust your digital advertising strategy
  • Analyze customer feedback across channels and initiate appropriate response workflows
  • Review legal documents and prepare necessary compliance filings
  • Manage complex project timelines while adapting to real-time changes

Practical Implications for Businesses

The combination of RAG and Agents represents a significant leap forward in business process automation. Organizations can now build systems that don't just provide information but actually complete complex workflows with minimal human intervention.

However, this power comes with responsibility. As these systems become more capable, it's crucial to implement proper governance structures, ensuring that AI actions align with business objectives and ethical considerations.

Looking Ahead

As both RAG and Agent technologies continue to mature, we're likely to see increasingly sophisticated applications that blur the line between knowledge systems and autonomous actors. The key will be finding the right balance between automation and human oversight, ensuring that these powerful tools enhance rather than replace human decision-making.

The future of AI isn't just about smarter systems – it's about systems that can both understand and act upon that understanding in meaningful ways. RAG and Agents are just the beginning of this transformative journey.

Saturday, December 14, 2024

Getting Started with AI: A 15-Hour Learning Journey

Want to become AI-savvy in just two weeks? Here's a focused learning path that requires only about an hour a day. This guide explores four essential themes that will transform how you interact with AI:

  1. Learning from AI Experts: How to leverage AI podcasts to build your knowledge foundation
  2. The Art of Iteration: Mastering the technique of refining your prompts to get better results
  3. Trust but Verify: Developing critical thinking skills to verify AI-generated content
  4. Smart Summarization: Converting lengthy AI conversations into powerful, reusable prompts

Let's dive into these themes through practical exercises and real-world examples that will help you harness AI effectively in your daily life.

Week 1: Building Your Foundation (7.5 hours)

Deep Dive into AI Through Podcasts (2.5 hours)

Start your journey by listening to carefully selected podcasts during your commute or daily routine. I recommend you choose only one or  two for your regular listening pleasure:

Mastering Prompt Iteration (2 hours)

Spend time practicing with AI chatbots, focusing on refining your prompts. Here's a fun example:

Initial Prompt:

"Write about dogs"

Improved Iteration:

"Write a 300-word guide about choosing the right dog breed for apartment living, including considerations for size, energy level, and noise"

Final Iteration:

"Create a comprehensive guide for apartment dwellers considering a dog. Include:

  • Top 5 breeds suited for apartment living
  • Exercise requirements for each breed
  • Noise levels and training tips
  • Space considerations
  • Estimated monthly costs

Format this as a practical guide with clear headings and bullet points"

Summary Iteration:

I often ask the chatbot to provide an improved prompt based upon the contents of the session.

  • "Please rewrite the prompts within this session into a single well-engineered prompt"

Asking the AI chatbot to rewrite your prompt really helps in deepening your understanding of prompt engineering.

And to make things interesting I sometimes ask the chatbot to rewrite the response with different literacy levels.

  • "Please rewrite this response for a grade five literacy level"
  • "Please rewrite this response for a PhD literacy level"
I actually find the response for the grade eight literacy level more interesting than the PhD level.

Using the different AI chatbots (3 hours)

There are many emerging AI chatbots, build some prompts within each. Experiment with different AI chatbots: Play, get curious, ask the AI bot to rewrite your prompt, try the rewrites against all these different chatbots, compare and contrast their responses.

  • ChatGPT: Excellent for creative writing and coding
  • Claude: Strong at analysis and detailed explanations
  • Gemini: Particularly good with multimodal tasks
  • Perplexity: Specialized in real-time information retrieval and citation

Week 2: Advanced Techniques (7.5 hours)

Verification Strategies (4 hours)

Learn to verify AI outputs effectively with these examples:

When you have Historical Facts:

"You mentioned the Wright brothers' first flight was in 1903. Can you:

  1. Provide specific sources for this date
  2. Break down the key events of that day
  3. Highlight any details you're uncertain about"

When you asked for Technical Advice:

"You've suggested this Python code solution. Can you:

  1. Explain why each line is necessary
  2. Identify potential edge cases
  3. Compare it with alternative approaches"
When you wanted Financial Analysis:

"You've provided a financial forecast for my small business. Can you:
  1. Explain the key assumptions behind your projections
  2. Identify potential economic factors that could impact these numbers
  3. Compare this forecast with industry benchmarks
  4. Highlight any areas where you have limited data or uncertainty
  5. Suggest additional data points that could improve the accuracy of this analysis"
These  are three examples of verification for your AI outputs. It is always a good idea to request verification as it reduces the AI hallucinations and increases your knowledge of the topic being discussed.

Work through the sessions from last week and write prompts to verify the information in an AI output. Spend a few hours creating verification prompts, ask the AI to write these for you. Improve upon the verification prompts, iterate.

An AI hallucination occurs when an artificial intelligence generates information that appears plausible but is factually incorrect or nonsensical. 

Session Summarization (3.5 hours)

Master the art of creating comprehensive prompts from AI sessions. Here's an example:

Original Conversation:

  • Human: "How can I improve my public speaking?"
  • AI: [Provides tips about preparation]
  • Human: "What about handling nervousness?"
  • AI: [Shares anxiety management techniques]
  • Human: "How should I structure my speech?"
  • AI: [Explains speech organization]

Summarized into Single New Prompt:

"Create a comprehensive public speaking guide for beginners that covers:

  1. Essential preparation steps
  2. Anxiety management techniques
  3. Speech structure and organization
  4. Delivery tips
  5. Common pitfalls to avoid

Include specific examples for each section and actionable steps for implementation"

Practical Exercise Examples

Try these exercises during your learning journey:

1. Content Creation:

  1.    Ask AI to write a blog post, then iterate three times, each time making it more specific
  2.    Example progression:
    •    "Write about healthy eating"
    •    "Write about healthy eating for busy professionals"
    •    "Create a 7-day meal prep guide for busy professionals who have only 30 minutes for dinner"

2. Problem Solving:

  •    Start with a complex problem like home organization
  •    Break it into smaller tasks
  •    Ask AI to verify the feasibility of each step
  •    Create a final, comprehensive action plan

Reminder: Ask AI to summarize a session and all its progressive steps into a single new prompt. Use this prompt in the different chatbots.

Key Takeaways

After completing this learning path, you'll have:

  • A solid understanding of current AI capabilities and limitations
  • Practical experience in prompt engineering
  • The ability to verify and validate AI outputs
  • Skills to maintain efficient AI conversations

Remember: Success with AI tools comes from systematic practice and refinement. Start with simple queries and gradually increase complexity as you become more comfortable with the interaction patterns.

Pro Tip: Keep a "prompt journal" documenting your most effective prompts and the situations where they worked best. This will help you develop your own library of reliable AI interaction strategies.

Tuesday, December 10, 2024

Finding Growth in the Gaps: How Career Breaks Fuel My Tech Journey

As a technology professional, I've discovered an unexpected rhythm in my career - one that turns the spaces between projects into powerful catalysts for growth. Every successful project completion brings not just a sense of accomplishment, but also a valuable gift: a few months of dedicated learning time. These self-directed sabbaticals, occurring naturally in my three-year career cycles, have become essential periods of exploration and reinvention.

Denis Hassabis and Hannah Fry
My current sabbatical feels particularly significant as I navigate the transformative world of Artificial Intelligence. Building upon my foundation in Machine Learning and data science, I've immersed myself in the AI landscape over the past two months. After exploring numerous AI podcasts, I've found two standout sources that consistently deliver valuable insights: "Google DeepMind: The Podcast" for cutting-edge AI research and developments, and "The Artificial Intelligence Show" by Marketing AI Institute for practical business applications.

This deep dive has also included extensive hands-on experimentation with leading Large Language Models (LLMs). Through countless hours working with ChatGPT, Claude, Gemini, and Perplexity, I've developed a nuanced understanding of each platform's strengths and refined my prompt engineering expertise. This practical experience has been invaluable in understanding the real-world capabilities and limitations of current AI technology.

This intensive learning period has already yielded tangible results. I've developed a comprehensive two-week learning module focused on AI fundamentals and practical applications, designed to help professionals enhance their productivity through AI tools. This resource embodies what I find most rewarding about these career interludes - the ability to synthesize new knowledge and share it with others who are eager to embrace technological advancement.

These deliberate pauses between opportunities aren't just breaks - they're investments in staying ahead of the technology curve. Each sabbatical allows me to emerge stronger, more knowledgeable, and better equipped to tackle the next challenge. In an industry that evolves at lightning speed, these learning periods have proven to be my secret weapon for sustained career growth and innovation.

Tuesday, October 29, 2024

Thank-you Keyin College

Another successful three years (well, three years and five months) working in an area I love; technology and adult learning. This was a great three years where I utilized my research and experience in integrating agile technologies and techniques into education. As a faculty we managed the complexity, and risk, of integrating an additional 40 students per semester into the program. And best of all, I taught over 200 students JavaScript, Node.js, SQL, PostgreSQL, Mongodb, GitHub, and how all these fit into FullStack development. It was definitely three years and five months well spent.

Three year successes

My career success seems to occur in three year cycles, I've written about this in the past. It would seem that my time with Keyin College follows the same pattern, I have just finished another three year cycle. Background to this three year theme (with focus on adult learning) can be found here;

  1. career success in three year cycles
  2. Increasing access to legal education

Gratitude

I like to share the success. I may have been the leader for a few things at Keyin College, but I couldn't have done it alone. These are the things where I am grateful for the staff at Keyin College;

Steve Taylor for encouraging me to participate in the early program design and for supporting me in becoming program head while we onboarded 60 students a semester.

Christa Mitchell for filling the role of career counselor while I was being onboarded to Keyin. So much to learn so little time. Also, for acting as my proxy when important items needed to be escalated to senior management. I know I can be demanding as an employee, I believe we should all be heard and an organization never knows where their diamond on the beach may be.

Sushanta for being good at want you do and your follow-through is exceptional!

Eric Bailey for helping me understand the overall technical architecture of Keyin's network, infrastructure, and approach to IT. Couldn't have made sense of things without you!

Maurice for being 5 hours older than me and providing the wisdom I needed to understand Keyin College (Or understand the best I could).

Roy Brushett and Will Durocher for those brainstorming sessions four years ago where we covered many important topics that defined the overall FullStack curriculum. Where Agile methodologies fit with Software Engineering and with Pedagogy, our use of O'Reilly for text books (saving students and Keyin considerable $$), the programming languages to be used, where source management (GitHub) fit within our Agile approach, Which board (Kanban or Scrum or both) we would use, etc. So good, great to be involved with the initial design and too see it through to teaching many courses as the whole program stabilized.

All the Faculty for allowing me the privilege of being in your service. We reduced risk and successfully went from 18 students to over 50 in one semester. Job well done. Thanks for your support. It was an honour to be your program head.

And of course, the Students! Honestly, I learn more than I teach. Your questions, your feedback, all our 1 on 1's, the tutorials, and encouraging me to develop workshops on an introduction to and intermediate github. Students are the best! We wouldn't be here without you!

Next Steps

I am being drawn back into the technical realm. Give teaching a break again for a while, focus on shipping software and building a team to exceed expectations. I am good at this so why not lean heavily into a startup, manage / mentor a team, leverage my knowledge of big data, database design and administration, draw upon the family (and personal) history of engineering, or even consider a director or executive position (if someone will have me, I'm so much better in a supporting role).

If you need someone with my background or experience please reach out. Thanks again Keyin College!

Tuesday, May 14, 2024

Blogging in a time of AI

I am renewing my frequency in blog posting. This will come after an almost 5 year break from blogging. I am returning because I am back to working on open-source software, educational projects, and the digitization of oceans. And I know I have a learning and success journey to share. 

Some Background

I started blogging twenty years back. Yes, 20 years. I was an early adopter and I was involved with technology startups and how the internet was influencing education. At this time I was a big believer that blogging was about content creation. Adding to the collective of the internet by adding meaningful and descriptive content rather than only being a consumer of content. To date, I have published over 500 posts to my assorted blogs. Most of this work was in the first 10 years of my blogging. I essentially posted once a month for the first 10 years. I did have a year where I posted over 100 times. As a summary this is how I posted over the last 20 years.

2004 to 2014 - I published 420 blog posts with good readership. I had a year where I posted more than 100 times. I set this as a goal / experiment to see if I could post twice a week. I did this while holding full-time work, which meant many early mornings and late evenings writing. I learned a whole bunch and my writing skills improved. 

The themes for these first ten years was mostly; 

  • technology leadership for startups, 
  • hard-core technology and approaches, 
  • innovative and emerging education, 
  • and the intersection of these three.

2015 to 2024 - I published only 41 blog posts during this 10 years, and nothing in the past three years. Honestly, I was distracted by raising my family of three and doing a whole bunch of life living. Not so focused on work and career advancement. 

The themes for the last ten years have mostly been; 

  • integrating with technology community in St. John's NL (I moved), 
  • continue my work on digital badges and micro-credentials,
  • development of an ocean data startup (still a work in progress),
  • and working the idea of a reference architecture for the digitization of oceans.

Most exciting of all this is more than 1/2 million views during the 20 years and at some points having over 2500 weekly views. What have I learned from all this blogging? Mostly, that having to write and publish openly to the internet helps the overall community knowledge and it helps me learn more deeply in these chosen subjects.

Next Steps

Again, I will use blogging as my cognitive gymnasium. My subjects themes haven't changed and I will focus upon two main subject areas and continue with updates to this critical technology blog;

  1. Education technology, Heutagogy, and the self-directed learner.
  2. Many things related to, and in support of, the authoring a reference architecture for the digitization of oceans.
  3. And my continued musings about technology through my gen X view of the world.

With all my work and R&D efforts I will learn a bunch of stuff and apply this to the real world through the successful projects I am a part of. I will reflect upon these successes and all that I have learned I will create content that can provide further learning for those around me. And hopefully they will also be entertaining reads.

Collaborating with your AI partner

Blogging has changed for me. There has been a lot of technical and social change since I did most of my blog posting over a decade ago. I had a few focused subjects I was very passionate about, and I wrote about them often. I wrote unincumbered for I would consider myself an early adopter and there was less people publishing to the internet in my chosen subjects. Today there is much more content covering these subjects. And all this content comes in photos, videos, audio, and written articles. Artificial Intelligence is doing a great job of creating and summarizing the content which addresses the complex audience needs and their questions and prompts. Content creation has changed. For a human content creator I believe our work needs to be more intelligent, critical, and creative. Content creators in a time of AI need to do what the AI cannot; daydream, reflect on unrelated subjects, see unlikely connections, be critical, add meaning, create new content that falls between the generated content, fact check and confirm, and add more human intelligence.

How will my blog writing process change? Um, it already has...

I must reflect and draw upon my mastery and do my best to add the new content that AI cannot... AI needs our creativity because it has already parsed the published body of human knowledge. For more insight on my approach, use your favorite large language model chatbot (ChatGPT, Gemini) with the prompt 'limitations of generative AI' followed by the prompt 'How would you suggest a human writer overcome these AI limitations'.

Step by Step my blogging will partner with AI and follow this basic approach;

  1. Capture ideas for new posts, be verbose, be imaginative, think about context
  2. Put these ideas to incomplete blog posts, work ideas for days, for weeks...
  3. Read extensively, add to the understanding of any specific idea
  4. Keep references, cut and paste to the bottom of the related incomplete posts
  5. Prompt AI with phrases from the idea generation
  6. Take blocks of text from written ideas and push them into generative AI, be critical, harvest what you can.
  7. Take the written blog post and ask AI for a rewrite. Change your audience. be critical, harvest what you can.
  8. Try and see, try and write, what AI cannot... add to the body of human knowledge.
  9. Add story telling to improve the overall post
  10. Find pictures to support the writing, format for engagement. Use AI to generate images from passages of text taken from the blog post.
  11. Format, edit, improve, repeat. Be bold... Publish.
  12. Use AI to improve the quality of the writing... Publish again.
  13. Rest, reflect, improve... Publish again.
  14. Yes, I am an advocate to publish before writing is perfect. Publish and then make improvements over the days and weeks that follow. Once the post is considered finished finished... promote it on social media.
  15. Identify what is most important about the post and rewrite for the LinkedIn business audience. Publish to LinkedIn.
  16. Repeat...


Friday, April 09, 2021

The Leanstack Way

The Oceans of Data Lab is honored to be a part of PropelICT's startup accelerator. We had our kickoff meeting a few days ago and the current focus is on learning the Leanstack methodology and using lean canvas to tease out ALL the important details for success. The inline supporting learning modules that are available through leanstack are very helpful. 


The Lean Canvas - numbers indicate the order of completion

I feel fortunate that I have been familiar with Agile and Lean approaches for over 20 years. I've got two favorite sayings I use when running software teams, and I like to think I run many aspects of my life with an Agile / Lean mindset.

  1. Ship and ship often (deliver new releases as often and frequently as possible)
  2. Fail and fail often (take risks, innovate, don't apologize, keep moving), success comes from failure.

For me the use of Lean in startups all began with Eric Ries when I watched a YouTube interview of Eric conducted a decade ago. This interview became a part of a 2011 blog post where I describe lean approaches within the Director of Technology role. Since this time I have revisited the works of Eric Ries every few years, he has a lot of useful insights to lean startups. One of my all time favorites in the talk he gave at Google 10 years back.


Google Talk: Eric Ries and the Lean Startup


Thursday, April 08, 2021

ODL Newsletter - March 2021

The Oceans of Data Lab (ODL) monthly newsletter is also finding its footing. It is still going to include monthly updates to the progress we make AND it will start with a few articles of interest within the data labs technology world. I am discovering so many interesting technologies and approaches within the data realm. I'm going to fold my 30 years of data experience into why I believe these are of interest to those working with large amounts of data.

Apache Data Lab

The Apache data lab that comes from the same organization that has brought us so many of the important technologies over the years. And specifically, to think of all the big data technologies they have delivered in recent years... there are just to many to list. What I like most about the data lab is its ability to be deployed to the big three cloud hosting environments. Super smart given the storage and compute requirements for data projects shouldn't be the responsibility of Apache.

DataOps and the DataKitchen

DataOps is a fairly recent concept / term that is about seven years old... and it makes sense that it becomes a discipline in itself as it is not DevOps for Data, it is so much more. The DataKitchen looks to be doing some amazing work in this capacity and have published a good read to help get your head into this important and emerging technology space.

I'm another 3 people into working towards my 100 conversations. It is said that you need to have 100 conversations as you solidify your business / startup idea. So I managed to get another three conversations in. I know this isn't that many, but that's ok as this month was more about setting up technology and thinking about risk, revenue, and the escalator pitch for the startup. I still need to talk with people, and I need peoples help, always. If you know anyone who works with analyzing data or works for a business that has a growing interest in their data, I'd love to talk with them.

What has changed this month?

Over this month my thinking has broadened and become more focused on the needs of organizations and their data. No real pivot, but clarifying what the business will be. The changes fell into three main themes;

  • A broader interest in helping people with their data. The backstory to my career has always been the information technology around the data. For 30 years I have focused on managing, moving, and building software for the data. This will continue with the data lab. We are still interested in ocean data and a reference architecture for the digitization of oceans, these subjects will become part of the bigger data lab.
  • It's a Data Lab. It became very clear this month that what I was wanting to do is stand up and run a data lab. I had a great conversation with Graham Truax at Innovation Island and this identified the alignment with my accelerator pitch and the data lab concept. After I re-read my proposal (and subsequent acceptance) to the PropelICT accelerator I confirmed... the startup is focused on creating a data lab with related products and services.
  • Start with a services focus, rather than product. We need revenue and the data lab is not a small product with a near MVP that can generate revenue. There are a number of MVP's that could bring business value for our customers, but nothing with significant revenue possibility. So our focus needs to be on services where we can leverage the skills and knowledge of the founder and identify projects that align well with the overall vision for Oceans of Data Lab.

It's been a business and technology focused month

This really was a more technology focused month. It was getting all the infrastructure in place to have the lab, fetch some data, and display a basic analytics dashboard. So while setting things up, we weren't that focused on reaching out to potential customers.

What are the risks and assumptions?

We also thought about what are the business risks and what assumptions are we making that could work against our success. We are not going to get into these in detail, writing them down and publishing them helps attract attention and hopefully getting the feedback we need to reduce the risk and prove or disprove the assumptions. We are also focused on what can be a product rather than what is a service.

Assumptions

  1. Companies / Organizations will participate in a publish - subscribe business model for data sets
  2. The data lab concept for preparing data sets for publishing will become accepted by SMB 

Risks

  1. MVP doesn't generate enough revenue or provide business value
  2. Primary founder having knowledge, energy, or bandwidth to keep up the pace
  3. Finding skilled employees with deep understanding of data engineering
  4. High cost of cloud based infrastructure 

Where is the revenue?

  • The transactional costs in the publish and subscribe (every data set transaction earns money)
    • this is definitely my riskiest assumption
  • SMB pay for services in preparing the data sets for themselves and the marketplace.
    • does the rise of the data engineer role show a willingness to pay for data preparation

What do we consider our Escalator Pitch?

These are early times and we don't yet have a story to tell. Gak! The escalator pitch is hard, and we really don't know what we are doing when it comes to an escalator pitch.

  • We help SMB realize new revenue possibilities from their existing data.
  • We reduce the cost of data preparation for their internal analysis and business intelligence.
  • We provide the services and technology to help you make sense of all your data. 
  • We make it easy for you to see the value and opportunities based upon your unique business data.

Next Steps:

  • We need to focus on the customer. We need to find the customers and talk to them.
  • We need to reduce our risk and prove, or disprove, our assumptions.
  • We need a technical platform to host an Minimum Viable Product (MVP). I need to identify and prioritize a few MVP's.

If you find the Data Lab an interesting idea or have the need to bring greater value to your existing data, please feel free to contact me. We are building a business and we want to help you bring greater value from your data.

Tuesday, March 23, 2021

It's Alive! The Elastic Stack as our Data Lab

So much technical work, so little time! I finished my first three sprints toward standing up the data lab. Standing up infrastructure from scratch so you have clean new compute power is fun, and also a lot of work. Particularly when you include; doing it right, taking no short cuts, and making sure it is secure.

Sprint 0: Setup Ubuntu 20.04 Server with ELK stack.

This was mostly rehydrating virtual server infrastructure I hadn't used in 8.5 years. It needed an upgrade from all perspectives and had a completely new OS. I implemented the ELK stack and made a couple of security changes to lock it all down. I ran a few tests by setting up a couple of websites, getting the JSON confirmation from ElasticSearch, and called up the Kibana dashboard. Oooo... sweet success!

Sprint 1: Vulnerability Assessment. Security changes if required.

This evening I spent some time poking at the overall vulnerability of the server and with the ElasticSearch and Kibana services. I made a few additions and changes for further locking down the services and believe they are as secure as they can be for this first release. Very happy to feel reasonably confident about it's being locked down. Maybe, I'll get lucky and get some free PEN testing. ha.

Sprint 2: Identify and register some well aligned domain names.

I registered the following domain names, even considered buying one... it would have been too expensive. I'll implement the data lab on the oceansofdatalab.com site when it becomes closer to being a minimal viable product (MVP).

  • oceansofdatalab.com
  • oceansofdatalab.org
  • oceansofdatalab.net
  • oceansofdata.net
  • sevenseasofdata.com
  • sevenseasofdata.org


Thursday, March 18, 2021

ODE Data Lab has its technology footing

This month has become more about standing up technology than it has been talking to people about their ocean data needs. That's ok... if you are building a technology company, you need to build technology. Ocean of Data Endeavours (ODE) is about building and utilizing software towards making it super easy to work with data, large amounts of data.



The last 10 days have been about refreshing a server infrastructure I stood up 12 years ago for a number of other projects. What was left was a couple of simple websites, some domain hosting, and all the related mail server infrastructure. All of this needed a complete refresh to be brought up to date;

  1. Rebuild the server infrastructure to have more horsepower. - DONE
  2. Upgrade the Ubuntu OS from 10.04 (Lucid Lynx) to 20.04 (Focal Fossa). - DONE
  3. Rework all the domain aliases to remove dependency on a domain I no longer owned. - DONE
  4. Do some basic security work to the server. Mostly SSH focused. - DONE
  5. Create a new mail server, and do some mailbox maintenance. - DONE
  6. Install Apache2 httpd host. - DONE
  7. Configure Apache2 for a couple of web sites. - DONE
  8. Code some basic HTML to confirm the sites are working. - DONE
  9. Celebrate! http://endeavours.com/

So good to have all this done. The server will provide a strong foundation and is well prepared for the ELK stack and the first load of ocean data. So excited! 


Tuesday, March 16, 2021

An Important difference between DevOps and DataOps

Where DevOps is automation, technology, and delivery focused; DataOps is more customer focused. I like these descriptions from Wikipedia for DevOps and DataOps;

  • DevOps is a set of practices that combines software development (Dev) and IT operations (Ops). It aims to shorten the systems development life cycle and provide continuous delivery with high software quality. https://en.wikipedia.org/wiki/DevOps
  • DataOps is an automated, process-oriented methodology, used by analytic and data teams, to improve the quality and reduce the cycle time of data analytics. While DataOps began as a set of best practices, it has now matured to become a new and independent approach to data analytics. DataOps applies to the entire data lifecycle from data preparation to reporting, and recognizes the interconnected nature of the data analytics team and information technology operations. https://en.wikipedia.org/wiki/DataOps
The similarities between these two are many, particularly from a process and automation perspective. I see DevOps really focused on delivering quality software, and DataOps focused on delivering visualized data analytics to the customer.

Customer focused DataOps assists with Agility

Having a customer focused data analytics team fits well with an Agile approach. The data analytics team needs very involved customer analysts (or product owners). The customer analyst identifies the KPI's, models, or intelligences that need to be fulfilled. These become part of the backlog, and as new sprints are defined they become focused on the item(s) of analytic. A sprint can be built around a few analytics, then iterate around the items for a DataOps sprint;
  1. Where is the data? How do we get at it?
  2. How do we best move it? How often? What are the security or privacy issues?
  3. What needs to be cleansed or transformed? Is the data at the correct granularity?
  4. Do we already have any related data to improve the intelligence? Is this a new build or do we use / alter an existing pipeline?
  5. What models or analytics do we apply?
  6. How do we best visualize the data?
Not to say that DevOps can't fit well within Agile approaches, it can.... the backlog is more technically focused and fits into the sprint more from a continuous perspective than a customer perspective. (What DevOps features go into a sprint are often negotiated with the product owner). The focus of DataOps is in shipping features that fulfill a visualized analytic or more... The focus of DevOps is in CD / CI...


This approach worked well for us when working on a Business Intelligence project and our nine week sprints usually focused around 4 to 9 KPI's. The organization was in aerospace, they had many legacy data sources with new data sources coming online. As with many organizations, they were in a state of improvement and transformation. Fitting new cubes, representing KPI's, into sprints allowed us to show progress and success. The biggest challenge wasn't in the technical or delivery side of getting the data to the customer. The challenge came in developing a data team where every team member understood the process end-to-end and the efforts required during each step of the DataOps pipeline. Acquiring, cleansing, and transforming data takes as much effort and understanding as visualizing the data for the customer.

Saturday, March 13, 2021

For Contract Database Administrator

Do you require contracted database administration? Medium to small organizations using database technology to store corporate data definitely need database administration to care and feed for there database technologies. This care and feeding includes, and is not limited to;

  • Install and maintain database servers.
  • Optimize database security.
  • Build and maintain ETL pipelines.
  • Performance tuning of databases.
  • Storage optimization for databases.
  • Implement DataOps for up-to-date business analytics and its need for continuous data.
  • Install, upgrade, and manage database applications.
  • Create automation, and schedule, repeating database tasks.
  • Ensure recoverability of database systems.
If you require any, or all, of these database administration task we can help. With over 30 years of database experience complimented with a technology degree in database management we can work remotely to keep your databases healthy and reduce business risk. Part-time or full-time, reasonable rates.

Friday, March 12, 2021

Data Challenge Panel Session: The Power of data

I listened in on this presentation about the power of data. In particular, having Susan Hunt from Canada's Ocean Supercluster as one of the presenters. All the presenters had very valuable insights. It was an hour well spent!

The Session: The Power of data

Description: Come learn with us. Our panel guests are engaged in transformational data products and projects. Together, we’ll learn what opportunities they are creating when using the newest technology to exploit the power of data. And we’ll talk about careers. The opportunities are limitless.

Moderator: Cathy Simpson | Chief Executive Officer, TechImpact Panelists: Susan Hunt | Chief Technical Officer, Canada’s Ocean Supercluster Justin Kamerman | Chief Product Officer, Instnt Inc. Jason Lee | Partner, MNP Technology Solutions

Items for my follow-up:

A number of subjects sparked my interest from the presenters discussions. I believe these are the three that need follow-up from the current ODE perspective:

  • Building Models for Data, or transforming Data for Models
  • What is the business case for the datalab / data workbench?
  • What is the current state of DataOps?


Peter, we need to integrate data more easily.

I liked this article on Fundamental truths when it comes to innovators. I do think Elon Musk and Jeff Bezos have figured out how to be crazy successful in business. I do wish they were more philanthropic with the monies they have accumulated from their success. I digress...

I like the idea of a fundamental truth as a foundation for your business that doesn't change through time. And when I think about my commitment to build a data services business, I have now started to think about what would be the fundamental truths?

  1. We need to access and integrate data more easily.
  2. We need better ways to visualize, communicate, and understand, the data.
  3. We would like to reduce the compute and storage costs of data.
These are what I came up with through my review of my initial thinking of fundamental truths. I know there will be a few more and these three will be edited as my idea grows and gets greater footing.

Monday, March 08, 2021

Building the Data Lab Technology Stack

The idea of building a data lab is emerging from my ocean data conversations and how to best utilize my knowledge and skillset within this opportunity. In my mind, the service offering would be twofold;

  1.  Data Engineering / Software Development consulting and services with focus on ocean data. We will do the heavy lifting of extracting, cleansing, transforming, and loading your data. And then we will help with analysis and visualizing the data. We are comfortable working in both the open source and Microsoft technology stacks.
  2. Standing up (and data loading) the technology stack for the data lab. You are going to need to host all this compute power and storage somewhere. It could be on-premise. Most likely, it will be in the cloud. We can help with this also. We could build it in Azure, using the Microsoft technology stack. Or we could build it using an Open Source stack on top of Linux in any of the hosted environments of Azure, AWS, or Rackspace.

https://stock.adobe.com/
How do you build a low priced, large compute, technology stack to support data engineering efforts, implement a data lab, and showcase these new services capabilities. The low price is the key factor given the current startup state of this ocean data endeavour. Particularly, when you think of the cost of compute for processing and storing large amounts of data. I believe the the best way forward is as follows;

  1. Use open source where you can. Fortunately, many of the infrastructures, tools, frameworks, and programming languages for the data lab are open source.
  2. Automate the build so it can be built and torn down with ease. This would eliminate the need for the stack to always be running.
  3. Store the data at it's source, if possible. Fetch, and load, the data when you automatically rebuild the stack. Keep in mind this limits the amount of big data you can store locally, and loading large amounts of data can cumulatively take days. Be mindful of this. 

Note: this stack is to showcase the services capabilities. A full data lab would also need the ability to both persist and fetch data. It's going to take some time to build the data lab!

The Data Lab Technology Stack

The deployment of this technology stack will use open source wherever possible running on a Linux (Ubuntu) Server hosted at Rackspace. The rational for these decisions are;

  • Little to No licensing costs 
  • Strong familiarity with Rackspace as hosting company
  • Existing domain name (endeavours.com) hosted with Rackspace
  • Extensive experience with Ubuntu Linux in a hosted environment
  • Familiarity with deploying data intensive solutions using the ELK stack
  • Experience programming in Python

Note: The deployment of this technology stack will happen in phases, where each phase will complete with some basic tests to ensure the stack behaves as desired.

Phase 0

Phase 0 will be a basic ELK stack running on an Ubuntu Linux server hosted at Rackspace with access via the endeavours.com web domain. The use case for where the data comes from, how we transform it, the analysis, and visualization is still to be determined. This use case will be used for testing this first iteration of the newly stood up data lab. Exciting times!

Phase 1

During phase 1 we will add the Python programming language to the technology stack and use it for two purposes;

  • Apply a model to the data using Python.
  • Present the processed data to a web page for display.

Phase 2

During phase 2 we will add Kafka as an infrastructure resource, identify some additional data sources, and pre-process the data before it gets loaded into ElasticSearch.

Phase 3 and beyond

Investigate the Apache Data lab stack, add Spark to our lab, add a data workbench...

Friday, March 05, 2021

ODE Newsletter - February 2021

I'm 7 people into working towards my 100 conversations. It is said that you need to have 100 conversations as you solidify your business / startup idea. So this is where I am, seven conversation in. If you know anyone who works with ocean data or works for a business that has an interest in oceans, I'd love to talk with them.

Given the time restraints of being deep into a large data / database migration project, I consider February has been a good month for conversations. It provided me a good view into the horizon of ocean data. I followed the conversations that were presented to me without me directing the focus. For this is the first month, and I have yet to gain clarity of the gaps of where I need more information. This makes sense given I am at the beginning and don't know what I don't know. Now that it is the end of February I have identified the need to talk with customers of ocean data. This could become a focus for March. The conversations for February unfolded in the following order, with the following summaries and highlights;

PropelICT (https://www.propelict.com/)

I reached out to a past co-worker in a leadership position within PropelICT. PropelICT is an Atlantic Canada e-accelerator for tech startups. The conversation was very encouraging and initiated my application to their April cohort. Looking forward to their support in the coming months (and years).

Highlights: 

  • The idea of 100 conversations.
  • My first suggested conversation contact. 
  • Being a candidate for their e-accelerator.

eOceans (https://www.eoceans.co/)

I spoke with one of the principals of eOceans. Time very well spent, Thank-you! So many details to be digested from this conversation. This organization clearly understands ocean data and where it intersects with social media! A bulleted list seems the best to call out the highlights;

  • There are many open standards and organizations working in this space. The data standards seem to be "standardizing" and there are many organizations working toward bringing the data standards together. More open organizations are contributing than the closed proprietary types. CIOOS is the standout for Canada. EU and US are much further down the standards and open data path than Canada.
  • Both ends [(data storage and end-points (IoT)] of the data collection are well serviced with lots of business and startup activity. It's the middle were the greater opportunity exists. It's with the data integration with consideration for all the standards and granularity. "It would be nice to dust off a 10 year old data set and be able to easily use it".
  • Working with ocean data initiatives is very project based and finding the revenue sources / the business model for an open reference architecture for the digitization of oceans could prove difficult.

Highlights: 

  • Many open organizations already working in the ocean data space. 
  • The business side of what you are exploring (reference architecture) may be difficult, so much work is project based and gov't funded. A reference architecture seems like an NGO or consortium kind of thing.
  • Middle ground of software and data integration could be a big need given my skillset.

Mentorship

Super fortunate to reconnect with an older friend who has loads of experience; small devices, programming, data, startups to a favorable exit, machine learning, etc... many skills that align well with what I am doing. And on top of all this, I really enjoy the meandering conversations we share!

The one area where there is a strong overlap towards my ocean data focus and the mentors previous experience with the integration of data. And yes he confirmed, integrating data from different devices to a common standard is a lot of work for creating a single view into a broad data realm.

Highlight: He agreed to provide me mentorship within this endeavour. So great!

New Brunswick Ocean Strategy: Our Opportunities in the Blue Economy

This was an excellent online conference put together by the Ocean Supercluster. What I did most was listen, and a good thing too... I have so much to learn. I really liked the breakout sessions where there was more individual participation. Some names, and acronyms are becoming more familiar too me. 

Highlight: A small list of contacts I could reach out to. All good!

TechNL (https://www.technl.ca/)

I spoke with one of the leaders in TechNL and we talked about what I am wanting to do with data, in particular, ocean data. The conversation pointed towards two relevant contacts;

Highlight: That if I am going to be successful in this endeavour I am going to need partners. The time required for setting up an organization isn't the best place for me to be focusing my time at this stage of the startup. And given the nature of this startup needing to work in the open, the partnership route may be the best way to go...

Canadian Integrated Ocean Observing System (https://cioos.ca/)

So fortunate to have the attention of two CIOOS employees! They were so gracious a provided a broad and deep amount of information regarding the state of ocean data. Super helpful! CIOOS clearly knows the data. The best way to summarize my conversation is by including the important questions and there answers;

With ocean data where is the greatest pain?

Resources as in financial and skills / knowledge.

At the more general project level; governance and the people who know how to organize and stewardship data through its lifecycle. This is more a reference to the industry in general... it's a project issue. And having the ability to integrate with a project that happened years ago...

Do open data standards have an influence?

Absolutely! There are many references to open data. Most of what we deal with are open.

How easily integrated are the existing data sets?

It’s getting better. It can be difficult to get an older data set and want to integrate it. These older sets often lack the granularity or metadata that makes it easier to ingest. There is a definite need here at a project level. Developing an expertise here could become a strong business.

Most initiatives within this space are project based. Which makes it difficult for longer initiatives that have some data sustainability. Rarely are there long term funding initiatives.

Highlights: 

  • So many acronyms, references and URLs. The CIOOS folks provided me many references all pointing in the right direction. Reference to some of the ISO standards. 
  • The need for better stewardship of data so as data ages it still has usefulness.

Pisces Research Project Management (https://piscesrpm.com/)

Another fortunate conversation with a person deep into ocean data and with the added bonus of being very technical. This was a contact I harvested from the New Brunswick Ocean Strategy Conference. There are may topics I could summarize from this outstanding conversation, much of the information confirmed things I discovered from the previous conversations described above. This is good!

I did pitch my idea about mooring buoys as a fixed points of data collection, and having these buoys like the personalized weather stations that have become so popular. This employee loved the idea.

The exciting part of this conversation was the discussion of the technical stack used within the open data within the oceans sector. It was good to add this to the knowledge I had of the proprietary technical stack used when I was managing the software engineering dept. at Provincial Aerospace.

What is the most common tech stack for Ocean Data?

This person has extensive experience working with Government Organizations and Academics. From what they have seen the most common, and emerging, technology stack includes;

    • Python
    • Assorted data storage approaches. Often NOT an RDBMS.
    • QGIS is common.

These are the tools he finds most effective and common. Using QGIS pushes you into the geo representation of data. Much ocean data requires different kinds of models, more 3d, more oceans… not necessarily geographic, etc.

The ability to prove models with real data is the biggest need from a technical perspective. This is why python has such good traction. It is easy for non-programmers and also rich enough for programmers. A good language for data, and useful across the technical skills working with data.

NetCDF is the most common data-store. Also CSV and proprietary data storage. Remember data people are mostly not programmers or overly technical.

Also take a look at CKAN (https://ckan.org/)

What are people looking for from a technical perspective?

    • Proving models with real data.
    • Integrating data

Highlights: 

  • A deep discussion about the technical stack. The preferred programming languages, data storage, integration approaches, and technical issues.
  • Confirmation that integrating data and proving models is an area of software development opportunity.

Lessons Learned

  1. A reference architecture for the digitization of oceans is not enough to hang a startup or business upon at this time! Where I do believe it is still a good idea that will form through time. There is so much work already going on for a common open architecture that another doesn't need to be started. I truly believe a reference architecture will emerge, it is a; when it will happen, not if it will happen.
  2. There is a big need for technical and software development skills and knowledge in the data engineering space of ocean data. I believe the opportunity exists for a software development / data engineering consulting firm with the specialty of ocean data.
  3. The idea of an anchored (or fixed) buoy for ocean data collection is very compelling too me. Kind of like the personal weather station but as a fixed mooring buoy. Anyone who has a mooring buoy could replace it with the data buoy, and have real-time data about the conditions at the buoy in preparation for mooring.

Next Steps

  1. March will be the month of broadening my reach. I need to talk with a broader section of people working in the oceans space. I need to find potential customers for the processing and software development in, and around, ocean data. 
  2. I need to start building software tools for the processing of ocean data. I need a reference technology stack showcasing our abilities to work with data.
  3. I need to start developing an elevator pitch for the ocean data software consulting firm. I need customers and revenue to get the real feedback to focus the business mission.

Sunday, February 21, 2021

The Beginning of Ocean Data Endeavours (ODE)

Thirty-nine months ago I started on a deep dive into developing a reference architecture for the digitization of oceans. The idea of developing this reference architecture was initiated by the Canadian Government awarding Atlantic Canada with the Ocean Super Cluster initiative and all my recent work with leading the software engineering group at Provincial Aerospace. My writing and research into ocean data took me to the point where I needed to deepen my understanding of a number of subjects, and I needed this deeper understanding before I could continue the writing and research (even though you could consider deepening understanding as research). I needed to have an intermediate understanding of what had come before and the current state of things with a reference architecture for the digitization of oceans. In particular, I needed to work directly with ocean data and the standards that influence its structure.

Over the last three years I have been lucky to work with Triware Technologies Inc., and together we have found projects that align with this need to deepen my understanding of all things digital and all things ocean. My recent project successes include;

OCIO Digital by Design - I was fortunate to be awarded the opportunity to be the data architect for the initial phase in digitizing the NL governments citizen facing portal. I remained on the project for the first 12 months through to the portals launch. Being on the design team to create the data tier and integrate with legacy data was a great achievement. And I deeply enjoyed using a scrum / jira approach with a multi-vendor, multi-disciplined team. We achieved a lot in a short period of time.

Lessons Learned - Agile, Scrum and Jira can scale well to a government organization with multiple scrum teams working toward an integrated solution.

Ocean Sector Search - We needed a way to index the Canadian Ocean Sector. So we built a search engine seeded by as many oceans related URLs as our analysts could gather. The technical architecture of this ocean specific search engine can be found in this previous post.

Lessons Learned - With reasonable technical effort Nutch can be configured and seeded to crawl a specific industry sector (in this case Canada Oceans Sector). The Nutch crawl harvested a significant number of pages (> 32000) that were then loaded into the ElasticSearch (ELK) stack while relevancy scoring each page along the way.

NLCHI - My work with the Newfoundland and Labrador Centre for Health Information (NLCHI) was a quick engagement to focus their requirements backlog into a few manageable sprints. I was super fortunate to help get an important project underway and gain insights into the concept of a customer focused data workbench for a specific subject domain.

Lessons Learned - The idea of a personal data workbench is very compelling when you consider the number of data sets already available in the oceans sector. And if we could fold in open and proprietary data sets, while honoring security and privacy we may be onto something...

Nalcor Energy Database Consolidation - So many databases, so little time. One of my favorite enterprise type projects is when the project pays for itself, over time, by the savings created by the projects downstream accomplishments. Not revenue generation, but operational expense savings. I believe one of the best KPI's for IT is not new systems implemented, but old systems retired.

Lessons Learned - an amazing amount of data can be moved with the correct use of tools, well built and managed ETL (pipelines), and a mindset of automation.

Argo Floats - 2018

NEXT STEPS

Over the last month I have revisited how to best develop my intermediate understanding of oceans data. After a number of conversations, with experts of oceans data, I believe my next steps are twofold; I need to focus on the existing standards for oceans data and I need to write some code to integrate some open oceans data sets.

I need to find opportunities to work directly with oceans data. If you are in the oceans sector, in any way, and you have the need of a very experienced data engineer, then I would love to help with your project. If you know of an oceans data project in needs of a data engineer, please forward on my credentials. Thanks to everyone for reading this far. And thank-you Triware for your ongoing career support!

Thursday, March 19, 2020

Digital by Design, Agility and Data Architecture

For 12 months, starting the summer of 2018, I was very fortunate to fill the data architect role for the Government of Newfoundland and Labrador's digital by design citizen facing web portal. An amazing team was brought together and we accomplished an amazing amount of work given the complexity of the environment we were all working. Kudos to the leadership team for seeding the ground and pulling together a diverse and effective group of people.

Being the oldest team member, with 35 years as a technology professional, I noticed a number of items and approaches that I consider the highlights of the project. I call out my 35 years experience because I know success doesn't always happen in a large group of people (with a team larger than 35). A group of strangers doesn't always come together when tasked to ship software on schedule and on budget. The cool part of this project is that the highlights were both technical and project management. In a nutshell, we came together using a scrum model of project management (hosted within JIRA) and architected a microservices technology stack using predominantly Microsoft technologies. The user experience design was exemplary and the software approach stayed aligned with the best of agile practices. We also used a scrum of scrums approach to manage the three distinct scrum teams.

What made this first year of a new project so effective?

The Agile Practices
The team was encouraged to use Agile approaches to successfully ship software. Thankfully, the commitment came from the most senior level and agile workshops were used to align the teams understanding and approach to agile. I consider these three agile practices what kept us all well aligned;
  1. We rigorously stayed with 3 week sprints. This was facilitated by the scrum of scrums group and kept us all focused on shipping working software.
  2. We embraced jira and stayed true to moving cards. It took a few sprints, as a whole we ended up having all the team members updating and moving cards. This, combined with morning standups, kept the team transparency high and important issues in the open.
  3. We always had demo days and retrospectives. This went a long way to keeping us focused and successful. All team members were encouraged to attend the other scrums demo days, this built excitement and kept us focused and moving.
Software Engineering Discipline
Developing software is as much art as it is science. Our team included many accomplished software engineers and this helped us implement features quickly and completely. Kudos are deserved by many on this project team, in particular, one of our technical leads (this is you, Phil) was hellbent and lead through example with two attributes of software engineering that are super important and sometimes missed;
  1. We refactored always, no excuses. As a group we were always learning, as implementing features is a relentless teacher. Improving upon our code base through refactoring kept the quality improving, and the bugs low. Even from a data architecture perspective, at the beginning of each sprint we refactored the data tier with the required data changes from the previous sprint. Data tiers often have different heart beats that the middle and user tiers as they are dependent on the legacy systems, which often have legacy heart beats. This is a blog post in itself...
  2. Automated testing. We automated whenever we could, we aspired to have automated tests with coverage to all our code. We got close by using frameworks and having a test first mind set. And don't underestimate how effective existing testing frameworks can be applied to the data tier.
Architecture was collaborative 
All architects were encouraged to contribute and discuss, we were always white-boarding and soliciting feedback. This kept the architecture strong and well understood throughout the team. And because we had a shared understanding of architecture the refactoring was reduced. All good...