14:47, 28th September 2021
Data Sources in Power BI Desktop
Power BI Desktop supports a wide range of data sources, from traditional databases and spreadsheets to cloud-based services and web resources. Connecting to these sources involves selecting the appropriate protocol and specifying details such as server addresses, URLs or file paths.
For scenarios requiring shared connection settings, PBIDS files offer a way to export and distribute connection configurations. These files use a structured JSON format to define protocols, addresses and optional parameters such as connection mode, with supported examples including Azure Analysis Services, SharePoint lists, SQL Server and web data sources.
When using PBIDS files, users must ensure compatibility with supported protocols and avoid including encrypted columns or unsupported features. The files can be created automatically through Power BI Desktop or edited manually in a text editor, offering flexibility in how connection details are defined and maintained. This approach facilitates collaboration and standardisation in data integration workflows.
14:03, 26th August 2021
Text Mining Node in SAS Model Studio on SAS Viya
The Text Mining node in SAS Model Studio on SAS Viya enables users to process unstructured data, such as free-form comments and reviews, and transform it into structured, quantitative representations through singular value decomposition, which can then be used as inputs for predictive modelling. When multiple variables carry a text role, the node defaults to the one with the greatest length, though users can override this by rejecting unwanted variables in the Data tab.
Configurable parsing options include part-of-speech tagging, noun group extraction, entity extraction and term stemming, whilst a minimum document threshold controls which terms are retained. From version Viya 4 2021.1.3 onwards, users can also upload custom lists such as stop lists and start lists.
The node generates up to 25 topic-based features, and in demonstrated comparisons a Decision Tree model incorporating those features outperformed one that did not, illustrating the potential value of extracting information from unstructured data. SAS Model Studio also supports automated pipeline creation, which detects text-role variables automatically and incorporates the Text Mining node into a full pipeline covering data preparation, model building, hyperparameter tuning and model selection. Results show that generated features frequently rank highly in variable importance across multiple model types.
14:02, 26th August 2021
Natural Language Processing: An Introduction
Natural language processing offers tools to extract meaningful insights from unstructured text data, enabling applications across a wide range of fields. Techniques such as tokenisation, sentiment analysis and entity recognition allow textual information to be transformed into structured formats that can be integrated with relational databases for predictive modelling and decision-making.
In healthcare, for example, analysing patient notes can reveal psychosocial factors influencing treatment outcomes. In legal contexts, automated summarisation of case documents aids in identifying key details. The technology also supports innovations such as chatbots and translation services, though its effectiveness relies heavily on the quality of input data.
As the field advances, the integration of text analytics with other data sources will become increasingly important for generating comprehensive insights. This is particularly relevant in domains where traditional data alone may not capture the full complexity of a problem.
14:02, 4th August 2021
Data Science Experience | SAS
The Data Science Experience highlights real-world applications of data science across industries, showcasing how professionals address complex challenges through innovative solutions. Examples include using integrated tools to enhance customer experiences in banking and developing strategies to ensure model reliability in digital transformation initiatives. These stories illustrate the importance of combining technical expertise with clear objectives to solve problems in healthcare, insurance and public sector contexts. Additional resources such as training programs, events and cloud-based analytics deployment options are available to support further exploration and skill development in the field.
09:01, 4th August 2021
NumFOCUS: A Nonprofit Supporting Open Code for Better Science
NumFOCUS supports open-source projects used by organisations ranging from major technology companies to research institutions, aiming to address complex challenges through collaborative development. It offers opportunities for community engagement, employment in open-source roles and ways for individuals and organisations to contribute financially or through sponsorship. The organisation also provides resources such as a newsletter, annual reports and a shop where proceeds support its initiatives, while maintaining a focus on fostering innovation and accessibility in scientific computing through its various programs and partnerships.
08:59, 27th May 2021
Apache ORC is a columnar storage format designed for Hadoop workloads, offering efficient data handling through features such as ACID transaction support, built-in indexes for rapid data retrieval and compatibility with complex data types including structs, lists and maps. It is maintained by the Apache Software Foundation, a non-profit organisation that oversees open-source projects under the Apache Licence, ensuring governance and privacy standards. The project provides documentation and tools for integration with various frameworks like Spark, Hive and Hadoop.
13:47, 16th May 2021
Machine Learning Operations
MLOps represents an emerging approach to managing the full lifecycle of machine learning development, aiming to integrate ML models into software engineering practices by unifying release cycles, enabling automated testing of models and data, applying agile methodologies and embedding ML components within continuous integration and delivery systems. It addresses challenges such as technical debt and emphasises cross-platform compatibility, while covering topics like workflow design, governance, deployment strategies and standardised processes for model development. The framework also includes tools for structuring infrastructure and defining governance practices, presented through various phases and principles to support the iterative and complex nature of ML-based software projects.
09:34, 13th May 2021
Business Process Model and Notation (BPMN) offers a standardised graphical method for organisations to visualise and communicate internal procedures, enhancing clarity in business operations and facilitating collaboration between entities. The BPMN specification, including versions such as 2.0 and supporting resources like quick guides and examples, provides a framework for consistent process representation. Certification programs, such as the OCEB 2 initiative, aim to validate expertise in enterprise BPMN through structured examinations, with credentials serving as a benchmark for professional competence in the field. Resources and frequently asked questions are available to support understanding and implementation of BPMN standards.
10:49, 12th May 2021
SAS User Group UK & Ireland
The Independent SAS Language Community is a London-based group with over 1,200 members, bringing together those with an interest in the SAS programming language across areas such as analytics, modelling and administration. The community holds regular meetups and has hosted more than 100 past events, covering topics ranging from mainframe modernisation and data analytics to the role of SAS in a world where languages such as R and Python are increasingly prevalent. The group is well regarded, holding a rating of 4.7 out of 5, and is organised by Andrew Ratcliffe among others. It actively seeks presenters, venues, volunteers, sponsors and ideas from those who wish to get involved and can also be found on LinkedIn.
10:48, 12th May 2021
Top YouTube Machine Learning Channels
KDnuggets recently identified the top 15 YouTube channels for machine learning based on a combination of views per video, subscriber count and video quantity, using a search criteria focused on relevance and activity over the past year. After excluding channels with fewer than 100,000 views or no updates in 12 months, one channel was omitted due to recent controversies, leaving a final list that highlights creators offering content ranging from foundational concepts to advanced applications in the field. The channels vary in focus, from educational tutorials to research insights, with descriptions provided where available to aid viewers in selecting content aligned with their learning goals.